community pool / model answer

LFM2.5-350M

chat · LiquidAI/LFM2.5-350M

549.9 tok/s · model median · 12 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
Radeon AI Pro R9700 32GB4-bit539.6 tok/smeasured87- GB4,096hipfire
RTX 3060 12GB8-bit574.7 tok/sreported1401.9 GB4,608llama.cpp
RX 5700 XT 8GB4-bit257.3 tok/sreported76- GB4,096hipfire
RX 6900 XT 16GB4-bit275.8 tok/sreported67- GB4,096hipfire
RX 7900 XTX 24GB4-bit398.4 tok/sreported57- GB4,096hipfire
RX 9070 XT 16GB4-bit479.2 tok/sreported78- GB4,096hipfire
Ryzen AI Max 395 103GB4-bit343.1 tok/sreported58- GB4,096hipfire

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.