community pool / model answer
Qwen3.6-27B
chat · unsloth/Qwen3.6-27B
27.5 tok/s · model median · 41 runs in the model record.
| Reference rig | Quantization | Median result | Basis | TTFT (ms) | Peak VRAM | Max context | Engines |
|---|---|---|---|---|---|---|---|
| Arc Pro B70 32GB ×2 | 4-bit | 42.1 tok/s | reported | - | - GB | 512 | llama.cpp |
| Arc Pro B70 32GB ×3 | 4-bit | 49.4 tok/s | measured | - | - GB | 2,048 | llama.cpp |
| Arc Pro B70 32GB ×4 | 4-bit | 39.2 tok/s | measured | - | - GB | 1,024 | llama.cpp |
| M3 Ultra 512GB | 4-bit | 26.3 tok/s | measured | 316 | - GB | 8,192 | llama.cpp |
| M3 Ultra 512GB | 8-bit | 18.5 tok/s | reported | 574 | - GB | 8,192 | lmstudio |
| M4 Max 48GB | 4-bit | 21.8 tok/s | reported | 1 | - GB | 4,096 | lmstudio |
| RTX 3090 24GB ×2 | 4-bit | 41.2 tok/s | measured | 346 | 24.3 GB | 131,712 | llama.cpp, vllm |
| RTX 3090 24GB ×2 | 8-bit | 45.5 tok/s | reported | 804 | 19 GB | 8,192 | llama.cpp |
| RTX 3090 Ti 24GB | 4-bit | 155.9 tok/s | reported | 68 | 18.4 GB | 4,096 | llama.cpp |
| RTX 4090 24GB | 3-bit | 49.4 tok/s | reported | 163 | - GB | 8,192 | llama.cpp |
| Ryzen AI Max 395 128GB | 4-bit | 12 tok/s | measured | - | 24 GB | 65,536 | hipfire, llama.cpp, ollama |
| Ryzen AI Max 395 128GB | 8-bit | 6.4 tok/s | measured | - | 39.3 GB | 32,768 | llama.cpp |
Context tested
The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.