community pool / model answer

Qwen3.8-27B-NVFP4-BF16-LMHead

chat · RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead

183.6 tok/s · model median · 5 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
GB10 Grace Blackwell 128GB4-bit25.3 tok/sreported240- GB131,072sglang
RTX 5090 32GB ×24-bit189.6 tok/sreported3657.4 GB262,144sglang
RTX PRO 6000 Blackwell 96GB4-bit296 tok/sreported47- GB262,144sglang

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.