community pool / model answer
Qwen3.6-27B-FP8
chat · Qwen/Qwen3.6-27B-FP8
72.5 tok/s · model median · 15 runs in the model record.
| Reference rig | Quantization | Median result | Basis | TTFT (ms) | Peak VRAM | Max context | Engines |
|---|---|---|---|---|---|---|---|
| Arc Pro B70 32GB ×2 | 8-bit | 20.1 tok/s | reported | - | - GB | 1,024 | vllm |
| GB10 Grace Blackwell 128GB | 8-bit | 9.5 tok/s | reported | - | - GB | 65,536 | sglang, vllm |
| RTX 3090 24GB ×2 | 8-bit | 46.3 tok/s | measured | 421 | 22 GB | 262,144 | vllm |
| RTX 3090 Ti 24GB ×2 | 8-bit | 70.3 tok/s | reported | 465 | 22 GB | 131,072 | vllm |
| RTX 5090 32GB ×2 | 8-bit | 179.9 tok/s | reported | 2,222 | 58.3 GB | 262,144 | vllm |
| RTX PRO 6000 Blackwell 96GB | 8-bit | 92.6 tok/s | measured | 222 | 96 GB | 131,072 | vllm |
| RTX PRO 6000 Blackwell 96GB | 16-bit | 96.2 tok/s | reported | 72 | - GB | 262,144 | vllm |
| RTX PRO 6000 Blackwell Workstation Edition 96GB | 8-bit | 24.8 tok/s | reported | 9,998 | 82.4 GB | 65,536 | vllm |
Context tested
The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.