community pool / model answer

Spark-X2.5-1.7B

chat · XHToken/Spark-X2.5-1.7B

186.3 tok/s · model median · 4 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
RTX 3060 12GB4-bit206.7 tok/sreported723.1 GB4,096llama.cpp
RTX 3060 12GB6-bit166.4 tok/sreported813.5 GB4,096llama.cpp
RTX 3060 12GB8-bit149.6 tok/sreported723.4 GB4,096llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.