community pool / model answer
Qwen3.8-27B-GGUF
chat · unsloth/Qwen3.8-27B-GGUF
38 tok/s · model median · 170 runs in the model record.
| Reference rig | Quantization | Median result | Basis | TTFT (ms) | Peak VRAM | Max context | Engines |
|---|---|---|---|---|---|---|---|
| A100 40GB | 4-bit | 41.1 tok/s | reported | 10,667 | 19 GB | 65,536 | llama.cpp |
| A100 40GB | 8-bit | 33.7 tok/s | reported | 10,194 | 30 GB | 65,536 | llama.cpp |
| Arc Pro B60 24GB | 4-bit | 21.2 tok/s | reported | 147 | - GB | 2,048 | llama.cpp |
| Arc Pro B60 24GB | 6-bit | 16 tok/s | reported | 796 | - GB | 2,048 | llama.cpp |
| Arc Pro B70 32GB | 4-bit | 31.2 tok/s | reported | 2,995 | - GB | 2,048 | llama.cpp |
| Arc Pro B70 32GB | 6-bit | 24.1 tok/s | reported | 354 | - GB | 2,048 | llama.cpp |
| Arc Pro B70 32GB | 8-bit | 16.1 tok/s | reported | 572 | - GB | 2,048 | llama.cpp |
| CMP 170HX 64GB | 4-bit | 37 tok/s | measured | 1,435 | - GB | 132,352 | llama.cpp |
| CMP 170HX 64GB | 8-bit | 23.5 tok/s | measured | 1,795 | - GB | 132,352 | llama.cpp |
| Multi-GPU 56GB ×2 | 8-bit | 36.7 tok/s | reported | 3,082 | 45.5 GB | 2,048 | llama.cpp |
| Radeon AI Pro R9700 32GB | 1-bit | 41.2 tok/s | measured | 455 | 7.3 GB | 32,768 | llama.cpp |
| Radeon AI Pro R9700 32GB | 4-bit | 39.5 tok/s | measured | 615 | 21.6 GB | 131,072 | llama.cpp, lucebox |
| Radeon AI Pro R9700 32GB | 6-bit | 21.6 tok/s | measured | 641 | 25.6 GB | 32,768 | llama.cpp |
| RTX 3060 12GB | 2-bit | 18.3 tok/s | measured | 150 | - GB | 32,768 | llama.cpp |
| RTX 3060 12GB ×2 | 3-bit | 15.5 tok/s | reported | - | - GB | 32,768 | llama.cpp |
| RTX 3060 12GB ×2 | 4-bit | 18.3 tok/s | measured | 1,506 | 19.4 GB | 131,072 | llama.cpp |
| RTX 3080 12GB | 2-bit | 44.4 tok/s | measured | 2,526 | 9.7 GB | 65,536 | llama.cpp |
| RTX 3090 24GB | 4-bit | 40.8 tok/s | reported | 68 | 17.6 GB | 32,768 | ollama |
| RTX 3090 24GB | 6-bit | 51.5 tok/s | reported | 206 | - GB | 2,048 | llama.cpp |
| RTX 3090 24GB ×2 | 4-bit | 73 tok/s | reported | 1,175 | - GB | 2,048 | bunn-llama |
| RTX 3090 24GB ×2 | 8-bit | 51.6 tok/s | measured | 230 | - GB | 262,144 | llama.cpp |
| RTX 4060 Ti 16GB ×2 | 4-bit | 53.6 tok/s | reported | - | - GB | 2,048 | llama.cpp |
| RTX 4070 12GB | 2-bit | 65.7 tok/s | reported | 227 | 9.3 GB | 2,048 | llama.cpp |
| RTX 5070 12GB | 2-bit | 46.3 tok/s | measured | 582 | 10.1 GB | 16,384 | llama.cpp |
| RTX 5090 32GB | 2-bit | 88.1 tok/s | reported | 82 | - GB | 2,048 | lmstudio |
| RTX 5090 32GB | 3-bit | 79.4 tok/s | measured | 382 | 17.3 GB | 131,072 | llama.cpp |
| RTX 5090 32GB | 4-bit | 91.3 tok/s | measured | 108 | 23.3 GB | 262,144 | llama.cpp, lmstudio |
| RTX 5090 32GB | 5-bit | 60.2 tok/s | measured | 424 | 23.2 GB | 131,072 | llama.cpp |
| RTX 5090 32GB | 6-bit | 60.2 tok/s | measured | 279 | 28.3 GB | 131,000 | llama.cpp, lmstudio |
| RTX PRO 6000 Blackwell 96GB | 2-bit | 100.7 tok/s | measured | - | - GB | 131,072 | llama.cpp |
| RTX PRO 6000 Blackwell 96GB | 3-bit | 88.9 tok/s | measured | - | - GB | 131,072 | llama.cpp |
| RTX PRO 6000 Blackwell 96GB | 4-bit | 79.3 tok/s | measured | - | - GB | 131,072 | llama.cpp |
| RTX PRO 6000 Blackwell 96GB | 5-bit | 67.8 tok/s | measured | - | - GB | 131,072 | llama.cpp |
| RTX PRO 6000 Blackwell 96GB | 6-bit | 57.5 tok/s | reported | - | - GB | 131,072 | llama.cpp |
| RTX PRO 6000 Blackwell 96GB | 8-bit | 50.2 tok/s | measured | 42,652 | - GB | 131,072 | llama.cpp |
| RTX PRO 6000 Blackwell 96GB | 16-bit | 61 tok/s | measured | 40,786 | - GB | 131,072 | llama.cpp |
| RX 7800 XT 16GB | 4-bit | 42 tok/s | reported | - | - GB | 121,000 | llama.cpp |
| RX 7900 XTX 24GB | 4-bit | 41.7 tok/s | measured | 384 | - GB | 180,224 | hipfire, llama.cpp |
| Ryzen AI Max 395 128GB | 4-bit | 26.7 tok/s | measured | 603 | 23.9 GB | 131,072 | llama.cpp, lucebox |
| Ryzen AI Max 395 128GB | 5-bit | 13.4 tok/s | reported | 642 | 23.1 GB | 4,096 | llama.cpp |
| Ryzen AI Max 395 128GB | 6-bit | 13.3 tok/s | reported | 730 | 25.7 GB | 4,096 | llama.cpp |
| Ryzen AI Max 395 128GB | 8-bit | 11.2 tok/s | reported | 746 | 31.1 GB | 4,096 | llama.cpp |
| Ryzen AI Max 395 64GB | 4-bit | 24.1 tok/s | measured | 4,332 | 18 GB | 8,192 | llama.cpp |
Context tested
The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.