community pool / model answer

Qwen3-8B

chat · Qwen/Qwen3-8B

52.3 tok/s · model median · 44 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
AMD Ryzen 7 7840HS4-bit12.3 tok/sreported-- GB512llama.cpp
GTX 1080 Ti 11GB4-bit50.7 tok/sreported-7.8 GB8,192llama.cpp
GTX 1080 Ti 11GB5-bit36.8 tok/sreported-8.5 GB8,192llama.cpp
Multi-GPU 24GB ×24-bit113.5 tok/sreported186- GB2,048llama.cpp
Radeon 890M 32GB4-bit15.9 tok/sreported-- GB4,096llama.cpp
RTX 3060 12GB4-bit43.7 tok/smeasured370- GB131,072llama.cpp
RTX 3060 Ti 8GB4-bit53 tok/smeasured1,2067.5 GB65,536llama.cpp
RTX 3090 Ti 24GB4-bit133 tok/sreported1237 GB8,192llama.cpp
RTX 5090 32GB4-bit227.1 tok/sreported185.6 GB8,192llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.