community pool / model answer

Qwen3.8-27B-MTP-NVFP4

chat · sakamakismile/Qwen3.8-27B-MTP-NVFP4

117.3 tok/s · model median · 8 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
RTX 5090 32GB4-bit117.7 tok/smeasured815- GB172,800vllm
RTX PRO 6000 Blackwell 96GB4-bit112.1 tok/sreported-- GB2,048vllm

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.