community pool / model answer

Qwen3-0.6B

chat · Qwen/Qwen3-0.6B

141.7 tok/s · model median · 23 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
AMD Ryzen 7 7840HS2-bit152.5 tok/sreported-- GB2,048llama.cpp
AMD Ryzen 7 7840HS4-bit121.1 tok/sreported-- GB2,048llama.cpp
AMD Ryzen 7 7840HS8-bit81.5 tok/sreported-- GB2,048llama.cpp
M4 Pro 24GB16-bit61 tok/sreported-1.4 GB2,048llama.cpp
RTX 3060 12GB8-bit310.4 tok/sreported3042.8 GB4,608llama.cpp
RTX 3060 Ti 8GB4-bit141.7 tok/smeasured545.2 GB131,072llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.