community pool / model answer

Qwen3.8-27B-FP8

chat · Qwen/Qwen3.8-27B-FP8

58.5 tok/s · model median · 5 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
Arc Pro B70 32GB ×28-bit58.5 tok/sreported108- GB256vllm
GB10 Grace Blackwell 128GB8-bit14.8 tok/sreported-- GB65,536vllm
RTX 5090 32GB8-bit16.8 tok/sreported14430.9 GB8,192vllm
RTX PRO 6000 Blackwell 96GB8-bit100.4 tok/sreported68- GB262,144vllm

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.