community pool / model answer

Spark-X2.5-4B

chat · XHToken/Spark-X2.5-4B

84.7 tok/s · model median · 5 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
RTX 3060 12GB3-bit76.1 tok/sreported1713.8 GB4,096llama.cpp
RTX 3060 12GB4-bit90.7 tok/sreported1613.5 GB4,096llama.cpp
RTX 3060 12GB5-bit89.2 tok/sreported1634 GB4,096llama.cpp
RTX 3060 12GB6-bit77.9 tok/sreported1754.6 GB4,096llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.