community pool / model answer
gemma-4-31B-it
chat · google/gemma-4-31B-it
24.8 tok/s · model median · 32 runs in the model record.
| Reference rig | Quantization | Median result | Basis | TTFT (ms) | Peak VRAM | Max context | Engines |
|---|---|---|---|---|---|---|---|
| Arc Pro B70 32GB | 4-bit | 16.4 tok/s | measured | - | - GB | 131,072 | llama.cpp |
| M1 Max 64GB | 8-bit | 8.4 tok/s | reported | 287 | - GB | 2,048 | ollama |
| Radeon AI Pro R9700 32GB | 4-bit | 28.8 tok/s | reported | - | - GB | 2,048 | llama.cpp |
| Radeon AI Pro R9700 32GB ×3 | 8-bit | 14.2 tok/s | measured | 1,089 | 45.7 GB | 794 | llama.cpp |
| Radeon AI Pro R9700 57GB | 4-bit | 3.3 tok/s | reported | 8,092 | - GB | 2,048 | llama.cpp |
| RTX 3090 24GB ×2 | 4-bit | 135.1 tok/s | reported | 71 | 45.1 GB | 2,048 | vllm |
| RTX 3090 24GB ×2 | 8-bit | 87.7 tok/s | reported | 110 | 21.6 GB | 8,192 | vllm |
| RTX 5090 32GB | 4-bit | 68.5 tok/s | reported | 115 | 17.8 GB | 8,192 | llama.cpp |
| RTX 5090 32GB | 6-bit | 46.5 tok/s | reported | 158 | 27.3 GB | 8,192 | llama.cpp |
| RTX 5090 32GB | 8-bit | 38 tok/s | reported | 250 | 64 GB | 2,048 | llama.cpp |
| RTX PRO 6000 Blackwell 96GB | 16-bit | 40 tok/s | reported | 55 | - GB | 2,048 | vllm |
| RTX PRO 6000 Blackwell 96GB ×4 | 16-bit | 98.2 tok/s | reported | 43 | - GB | 2,048 | vllm |
| RTX PRO 6000 Blackwell Workstation Edition 96GB | 8-bit | 58.6 tok/s | reported | 68 | - GB | 131,072 | vllm |
| RTX PRO 6000 Blackwell Workstation Edition 96GB | 16-bit | 47.4 tok/s | measured | 26,513 | 86.1 GB | 131,072 | vllm |
| Ryzen AI Max 395 128GB | 4-bit | 10.7 tok/s | reported | - | 25.4 GB | 16,384 | llama.cpp |
| Tesla P100-PCIE-16GB 16GB ×2 | 5-bit | 15.3 tok/s | reported | - | - GB | 2,048 | llama.cpp |
Context tested
The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.