community pool / model answer
Qwen3.6-27B
chat · Qwen/Qwen3.6-27B
43.7 tok/s · model median · 146 runs in the model record.
| Reference rig | Quantization | Median result | Basis | TTFT (ms) | Peak VRAM | Max context | Engines |
|---|---|---|---|---|---|---|---|
| Arc Pro B60 24GB ×2 | 8-bit | 9.5 tok/s | reported | 1,135 | - GB | 4,096 | llama.cpp |
| Arc Pro B70 32GB | 4-bit | 43.5 tok/s | measured | 1,156 | 17.5 GB | 131,072 | llama.cpp, vllm |
| Arc Pro B70 32GB | 5-bit | 26.6 tok/s | measured | 311 | - GB | 204,800 | llama.cpp |
| Arc Pro B70 32GB ×2 | 4-bit | 40.5 tok/s | reported | - | - GB | 1,024 | llama.cpp |
| Arc Pro B70 32GB ×2 | 8-bit | 25.7 tok/s | reported | - | - GB | 512 | llama.cpp |
| Arc Pro B70 32GB ×3 | 4-bit | 44.2 tok/s | measured | - | - GB | 1,024 | llama.cpp |
| Arc Pro B70 32GB ×4 | 4-bit | 33.1 tok/s | reported | - | - GB | 1,024 | llama.cpp |
| Arc Pro B70 32GB ×4 | 8-bit | 43.7 tok/s | reported | - | - GB | 4,096 | vllm |
| Arc Pro B70 32GB ×4 | 16-bit | 54.2 tok/s | reported | 79 | 17 GB | 262,144 | vllm |
| AMD Ryzen 7 7840HS | 4-bit | 3.5 tok/s | reported | - | - GB | 2,048 | llama.cpp |
| GB10 Grace Blackwell 128GB | 4-bit | 33.3 tok/s | measured | 632 | 117 GB | 262,144 | llama.cpp, vllm |
| M1 Max 64GB | 8-bit | 11.3 tok/s | reported | 3,973 | - GB | 787 | llama.cpp |
| M5 Max 128GB | 4-bit | 19 tok/s | reported | 174 | - GB | 8,192 | lmstudio |
| Multi-GPU 28GB ×2 | 4-bit | 21.9 tok/s | reported | 155 | - GB | 4,096 | llama.cpp |
| Radeon 8060S Graphics 103GB | 4-bit | 88.3 tok/s | reported | 337 | 18 GB | 4,096 | hipfire |
| Radeon 890M 32GB | 4-bit | 4.9 tok/s | reported | - | - GB | 4,096 | llama.cpp |
| Radeon AI Pro R9700 32GB | 4-bit | 278 tok/s | reported | 103 | 16.6 GB | 2,048 | hipfire |
| Radeon AI Pro R9700 32GB ×3 | 4-bit | 213.7 tok/s | reported | 63 | 16 GB | 4,096 | hipfire |
| Radeon AI Pro R9700 32GB ×3 | 8-bit | 17.8 tok/s | measured | 768 | 49.4 GB | 262,144 | llama.cpp |
| Radeon AI Pro R9700 34GB | 4-bit | 254.8 tok/s | reported | 111 | 17.5 GB | 4,096 | hipfire |
| Radeon AI Pro R9700 57GB | 4-bit | 3.6 tok/s | reported | 7,848 | - GB | 2,048 | llama.cpp |
| Radeon RX 7900 XTX 24GB | 4-bit | 28.4 tok/s | measured | 490 | - GB | 262,144 | llama.cpp |
| RTX 3060 12GB ×2 | 4-bit | 17.6 tok/s | reported | - | 18.2 GB | 128,000 | llama.cpp |
| RTX 3090 24GB | 4-bit | 26.2 tok/s | measured | 1,374 | 0 GB | 131,072 | llama.cpp |
| RTX 3090 24GB ×2 | 4-bit | 46.7 tok/s | reported | 125 | - GB | 2,048 | llama.cpp |
| RTX 3090 24GB ×2 | 6-bit | 22.7 tok/s | reported | 2,101 | - GB | 2,048 | llama.cpp |
| RTX 3090 24GB ×2 | 8-bit | 40.9 tok/s | reported | 2,101 | - GB | 2,048 | vllm |
| RTX 3090 Ti 24GB | 4-bit | 140.3 tok/s | reported | 710 | 22 GB | 1,024 | llama.cpp |
| RTX 4090 24GB | 4-bit | 44.2 tok/s | measured | 335 | - GB | 131,072 | llama.cpp |
| RTX 5060 Ti 16GB | 4-bit | 46 tok/s | reported | 453 | - GB | 2,048 | llama.cpp |
| RTX 5090 32GB | 3-bit | 73.5 tok/s | measured | 36 | 13.5 GB | 8,192 | llama.cpp |
| RTX 5090 32GB | 4-bit | 70.1 tok/s | measured | 37 | 20 GB | 243,040 | llama.cpp, vllm |
| RTX 5090 32GB | 8-bit | 20 tok/s | reported | 200 | 16 GB | 2,048 | ollama |
| RTX PRO 6000 Blackwell 96GB | 4-bit | 105.7 tok/s | measured | 1,150 | - GB | 15,907 | llama.cpp |
| RTX PRO 6000 Blackwell 96GB | 8-bit | 72.1 tok/s | measured | 1,203 | - GB | 15,907 | llama.cpp |
| RTX PRO 6000 Blackwell 96GB | 16-bit | 56.4 tok/s | reported | 99 | - GB | 2,048 | vllm |
| RTX PRO 6000 Blackwell 96GB ×2 | 16-bit | 99.6 tok/s | reported | 58 | - GB | 2,048 | vllm |
| RTX PRO 6000 Blackwell 96GB ×4 | 16-bit | 146.1 tok/s | reported | 51 | - GB | 2,048 | vllm |
| RX 7900 XT 20GB | 4-bit | 60.6 tok/s | reported | 85 | - GB | 2,048 | vllm |
| RX 7900 XTX 24GB | 4-bit | 118.2 tok/s | measured | 126 | 16.8 GB | 16,384 | hipfire, llama.cpp |
| Ryzen AI Max 395 128GB | 4-bit | 112 tok/s | reported | 290 | 16.8 GB | 2,048 | hipfire |
| Ryzen AI Max 395 128GB | 8-bit | 16.7 tok/s | reported | 1,926 | - GB | 787 | llama.cpp |
| Ryzen AI Max 395 64GB | 4-bit | 12.4 tok/s | measured | 4,572 | 17.1 GB | 10,240 | llama.cpp |
| Ryzen AI Max 395 64GB | 8-bit | 7.7 tok/s | reported | 4,059 | 27.8 GB | 8,192 | llama.cpp |
| Tesla P100 32GB ×2 | 4-bit | 47.4 tok/s | measured | 176 | - GB | 2,048 | llama.cpp |
| Tesla P100-PCIE-16GB 16GB | 3-bit | 7.5 tok/s | reported | - | - GB | 2,048 | llama.cpp |
| Tesla P100-PCIE-16GB 16GB ×2 | 5-bit | 17.9 tok/s | reported | - | - GB | 2,048 | llama.cpp |
Context tested
The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.