community pool / model answer
gemma-4-26B-A4B-it
chat · google/gemma-4-26B-A4B-it
60.6 tok/s · model median · 43 runs in the model record.
| Reference rig | Quantization | Median result | Basis | TTFT (ms) | Peak VRAM | Max context | Engines |
|---|---|---|---|---|---|---|---|
| Arc Pro B70 32GB | 8-bit | 43.3 tok/s | reported | 393 | - GB | 2,048 | llama.cpp |
| AMD Ryzen 7 8745H w/ Radeon 780M | 4-bit | 30.6 tok/s | reported | 1,534 | 4 GB | 256,000 | llama.cpp |
| M1 Max 64GB | 8-bit | 49 tok/s | reported | 462 | - GB | 2,048 | llama.cpp, ollama |
| M3 Ultra 512GB | 4-bit | 76.9 tok/s | reported | 279 | - GB | 2,048 | mlx |
| Radeon AI Pro R9700 32GB ×2 | 4-bit | 87.3 tok/s | reported | 375 | - GB | 794 | llama.cpp |
| Radeon AI Pro R9700 32GB ×2 | 8-bit | 99.3 tok/s | measured | - | - GB | 262,144 | llama.cpp |
| Radeon AI Pro R9700 32GB ×3 | 8-bit | 73.7 tok/s | measured | 298 | 35.4 GB | 794 | llama.cpp |
| RTX 3090 24GB ×2 | 8-bit | 154.1 tok/s | reported | 42 | 45.3 GB | 262,144 | vllm |
| RTX 4070 12GB | 4-bit | 20.4 tok/s | measured | 1,661 | - GB | 2,048 | vllm |
| RTX 5090 32GB | 8-bit | 282 tok/s | measured | 30 | - GB | 2,048 | vllm |
| RTX PRO 6000 Blackwell 96GB | 8-bit | 276.7 tok/s | reported | 28 | - GB | 8,192 | vllm |
| RX 570 4GB ×2 | 4-bit | 12.5 tok/s | measured | 760 | - GB | 2,048 | llama.cpp |
| Ryzen AI Max 395 128GB | 4-bit | 60.7 tok/s | measured | - | 21.5 GB | 32,768 | llama.cpp |
| Ryzen AI Max 395 128GB | 6-bit | 55.5 tok/s | reported | - | 23.8 GB | 16,384 | llama.cpp |
| Ryzen AI Max 395 128GB | 8-bit | 46.1 tok/s | reported | 746 | 27.8 GB | 16,384 | llama.cpp |
| Tesla P100-PCIE-16GB 16GB ×2 | 5-bit | 54.8 tok/s | reported | - | - GB | 2,048 | llama.cpp |
Context tested
The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.