community pool / model answer

Qwen3.8-27B

chat · Qwen/Qwen3.8-27B

47.8 tok/s · model median · 67 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
Arc Pro B70 32GB4-bit28.9 tok/smeasured157- GB131,072llama.cpp, vllm
GB10 Grace Blackwell 256GB16-bit8.2 tok/sreported283- GB2,048vllm
H100 NVL 94GB ×28-bit293.3 tok/sreported69- GB2,048sglang
Radeon AI Pro R9700 32GB4-bit285.5 tok/smeasured11522.5 GB98,304hipfire
Radeon AI Pro R9700 32GB ×28-bit46.8 tok/sreported-- GB262,144llama.cpp
Radeon AI Pro R9700 32GB ×34-bit28.4 tok/smeasured919- GB829llama.cpp
Radeon AI Pro R9700 32GB ×38-bit45.7 tok/smeasured90335 GB10,952llama.cpp, lucebox
RTX 3090 24GB4-bit113.2 tok/sreported9122.4 GB65,536vllm
RTX 3090 24GB6-bit53 tok/sreported208- GB2,048llama.cpp
RTX 3090 24GB ×24-bit72.8 tok/smeasured949- GB262,144vllm
RTX 5060 Ti 16GB ×24-bit63.2 tok/smeasured190- GB131,072llama.cpp, vllm
RTX 5090 32GB4-bit157.4 tok/smeasured28024 GB170,000llama.cpp, vllm
RTX PRO 6000 Blackwell 96GB16-bit64.7 tok/sreported92- GB262,144vllm
RX 7800 XT 16GB4-bit28 tok/sreported-- GB2,048llama.cpp
RX 7900 XTX 24GB4-bit122.5 tok/sreported43417.3 GB32,768hipfire, llama.cpp
Ryzen AI Max 395 128GB4-bit48.4 tok/smeasured4,73317.6 GB4,096hipfire, llama.cpp
Ryzen AI Max 395 128GB8-bit14.2 tok/smeasured2,090- GB829llama.cpp
Ryzen AI Max 395 64GB4-bit18.5 tok/smeasured4,37917.4 GB8,192llama.cpp
Ryzen AI Max 395 64GB6-bit17.3 tok/smeasured4,77225.7 GB8,192llama.cpp
Tesla V100 16GB ×44-bit391 tok/sreported34961 GB32,768vllm

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.