community pool / model answer
Qwen3.6-35B-A3B
chat · Qwen/Qwen3.6-35B-A3B
70.8 tok/s · model median · 136 runs in the model record.
| Reference rig | Quantization | Median result | Basis | TTFT (ms) | Peak VRAM | Max context | Engines |
|---|---|---|---|---|---|---|---|
| Arc Pro B60 24GB | 4-bit | 42.6 tok/s | reported | 590 | - GB | 2,048 | llama.cpp |
| Arc Pro B70 32GB | 4-bit | 70 tok/s | measured | 372 | 22 GB | 32,768 | llama.cpp, vllm |
| Arc Pro B70 32GB | 5-bit | 67 tok/s | measured | 69 | - GB | 262,144 | llama.cpp |
| Arc Pro B70 32GB ×2 | 4-bit | 73.9 tok/s | reported | - | - GB | 2,048 | llama.cpp |
| Arc Pro B70 32GB ×2 | 8-bit | 110.7 tok/s | reported | - | 31.2 GB | 65,536 | vllm |
| Arc Pro B70 32GB ×4 | 1-bit | 120.8 tok/s | reported | - | 29.8 GB | 65,536 | vllm |
| Arc Pro B70 32GB ×4 | 4-bit | 68.8 tok/s | reported | - | 60 GB | 262,144 | llama.cpp |
| Arc Pro B70 32GB ×4 | 8-bit | 99.8 tok/s | reported | 77 | 127.6 GB | 32,768 | vllm |
| Arc Pro B70 32GB ×4 | 16-bit | 102.5 tok/s | reported | 104 | 123 GB | 262,144 | vllm |
| Intel(R) Core(TM) Ultra 7 155H | 3-bit | 13.4 tok/s | reported | - | - GB | 2,048 | llama.cpp |
| GB10 Grace Blackwell 128GB | 4-bit | 64.5 tok/s | reported | 153 | 87 GB | 262,144 | vllm |
| GB10 Grace Blackwell 128GB | 6-bit | 56.1 tok/s | reported | - | - GB | 2,048 | llama.cpp |
| GB10 Grace Blackwell 128GB | 8-bit | 55.7 tok/s | measured | 206 | - GB | 262,144 | vllm |
| GTX 1060 6GB | 2-bit | 15 tok/s | reported | - | 5.9 GB | 76,800 | llama.cpp |
| GTX 1080 Ti 11GB | 2-bit | 21.6 tok/s | measured | - | 10.4 GB | 8,192 | llama.cpp |
| GTX 1080 Ti 11GB | 3-bit | 14.6 tok/s | reported | - | 9.6 GB | 8,192 | llama.cpp |
| GTX 1080 Ti 11GB | 4-bit | 10.3 tok/s | reported | - | 10.2 GB | 8,192 | llama.cpp |
| GTX 1080 Ti 11GB | 8-bit | 6.4 tok/s | reported | - | 10.9 GB | 8,192 | llama.cpp |
| M1 Max 64GB | 4-bit | 59.3 tok/s | reported | 45 | - GB | 2,048 | ollama |
| M3 Ultra 512GB | 4-bit | 31.9 tok/s | reported | 620 | - GB | 2,048 | mlx |
| M5 Max 128GB | 4-bit | 119.7 tok/s | reported | 37 | 19.9 GB | 8,192 | lmstudio, mlx |
| Multi-GPU 28GB ×2 | 4-bit | 85.6 tok/s | measured | 81 | - GB | 4,096 | llama.cpp |
| Radeon AI Pro R9700 32GB | 4-bit | 277.4 tok/s | measured | 520 | - GB | 4,096 | hipfire, llama.cpp |
| Radeon AI Pro R9700 32GB ×2 | 4-bit | 32.6 tok/s | reported | 135 | - GB | 32,768 | vllm |
| Radeon AI Pro R9700 32GB ×3 | 4-bit | 98.9 tok/s | measured | 171 | 32 GB | 787 | llama.cpp, vllm |
| Radeon AI Pro R9700 32GB ×3 | 8-bit | 79.2 tok/s | measured | 343 | 44.8 GB | 787 | llama.cpp |
| Radeon AI Pro R9700 34GB | 4-bit | 259.5 tok/s | reported | 77 | 27.8 GB | 4,096 | hipfire |
| Radeon AI Pro R9700 57GB | 4-bit | 33.2 tok/s | reported | 1,294 | - GB | 2,048 | llama.cpp |
| Radeon AI Pro R9700 57GB | 6-bit | 26.4 tok/s | reported | 1,568 | - GB | 2,048 | llama.cpp |
| Radeon AI Pro R9700 57GB | 8-bit | 24.4 tok/s | reported | 1,254 | - GB | 2,048 | llama.cpp |
| Radeon RX 9070 16GB | 4-bit | 34.2 tok/s | reported | 417 | 14.6 GB | 131,072 | llama.cpp |
| RTX 3050 8GB | 3-bit | 31 tok/s | reported | 139 | - GB | 28,000 | llama.cpp |
| RTX 3060 12GB | 3-bit | 27.9 tok/s | measured | 1,476 | 5.6 GB | 4,096 | llama.cpp |
| RTX 3060 12GB | 4-bit | 35 tok/s | reported | - | - GB | 131,072 | llama.cpp |
| RTX 3060 12GB ×2 | 4-bit | 67.6 tok/s | reported | - | 20.9 GB | 96,000 | llama.cpp |
| RTX 3070 Ti 8GB | 3-bit | 33.5 tok/s | reported | - | 7,949 GB | 262,144 | llama.cpp |
| RTX 3080 12GB | 4-bit | 35.4 tok/s | measured | - | 9.5 GB | 196,608 | llama.cpp |
| RTX 3090 24GB | 4-bit | 11 tok/s | reported | 647 | - GB | 2,048 | llama.cpp |
| RTX 3090 24GB ×2 | 4-bit | 155.6 tok/s | measured | 121 | 21.3 GB | 32,768 | vllm |
| RTX 3090 24GB ×2 | 6-bit | 137.9 tok/s | reported | 648 | - GB | 2,048 | llama.cpp |
| RTX 4070 12GB | 4-bit | 39.6 tok/s | measured | 1,476 | - GB | 2,048 | vllm |
| RTX 4090 24GB | 3-bit | 157.7 tok/s | reported | 175 | - GB | 8,192 | llama.cpp |
| RTX 5080 16GB | 3-bit | 150.6 tok/s | reported | - | 14.2 GB | 131,072 | llama.cpp |
| RTX 5090 32GB | 4-bit | 190 tok/s | measured | 54 | 26.7 GB | 262,144 | llama.cpp, ollama, vllm |
| RTX 5090 32GB | 6-bit | 219.8 tok/s | reported | 21 | 27.5 GB | 4,096 | llama.cpp |
| RTX PRO 6000 Blackwell Workstation Edition 96GB | 16-bit | 164.1 tok/s | reported | 3,449 | 85.3 GB | 65,536 | sglang |
| RX 6700 10GB | 3-bit | 16.2 tok/s | reported | 2,448 | 4 GB | 4,096 | llama.cpp |
| RX 7900 XTX 24GB | 4-bit | 175 tok/s | measured | 76 | 22.1 GB | 4,096 | hipfire |
| Ryzen AI Max 395 103GB | 4-bit | 174.9 tok/s | reported | 296 | - GB | 4,096 | hipfire |
| Ryzen AI Max 395 64GB | 4-bit | 62.4 tok/s | measured | 1,020 | 21.5 GB | 8,192 | llama.cpp |
| T4 16GB | 2-bit | 54.7 tok/s | reported | - | - GB | 2,048 | llama.cpp |
| Tesla P100 32GB ×2 | 4-bit | 78.1 tok/s | reported | 157 | - GB | 2,048 | llama.cpp |
| Tesla V100 32GB ×2 | 6-bit | 96.7 tok/s | reported | - | 60 GB | 65,536 | llama.cpp |
| Tesla V100 SXM2 32GB | 4-bit | 66.8 tok/s | reported | 40,000 | 28 GB | 96,000 | ollama |
| Ultra 7 155H 32GB | 3-bit | 8.7 tok/s | reported | - | - GB | 2,048 | llama.cpp |
Context tested
The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.