community pool / model answer

LFM2.5-8B-A1B-GGUF

chat · unsloth/LFM2.5-8B-A1B-GGUF

66.8 tok/s · model median · 6 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
AMD Ryzen 5 7640U4-bit28.4 tok/sreported432- GB2,048vllm
RTX 3050 8GB8-bit47.7 tok/sreported95- GB28,000llama.cpp
RX 570 4GB ×24-bit54.6 tok/sreported125- GB2,048llama.cpp
RX 7700 XT 12GB4-bit152.4 tok/sreported137- GB2,048vllm

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.