community pool / model answer

Qwen3.6-27B-FP8

chat · Qwen/Qwen3.6-27B-FP8

72.5 tok/s · model median · 15 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
Arc Pro B70 32GB ×28-bit20.1 tok/sreported-- GB1,024vllm
GB10 Grace Blackwell 128GB8-bit9.5 tok/sreported-- GB65,536sglang, vllm
RTX 3090 24GB ×28-bit46.3 tok/smeasured42122 GB262,144vllm
RTX 3090 Ti 24GB ×28-bit70.3 tok/sreported46522 GB131,072vllm
RTX 5090 32GB ×28-bit179.9 tok/sreported2,22258.3 GB262,144vllm
RTX PRO 6000 Blackwell 96GB8-bit92.6 tok/smeasured22296 GB131,072vllm
RTX PRO 6000 Blackwell 96GB16-bit96.2 tok/sreported72- GB262,144vllm
RTX PRO 6000 Blackwell Workstation Edition 96GB8-bit24.8 tok/sreported9,99882.4 GB65,536vllm

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.