community pool / model answer

gemma-4-31B-it-GGUF

chat · unsloth/gemma-4-31B-it-GGUF

15.4 tok/s · model median · 22 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
Minisforum UM790 Pro 64GB4-bit4 tok/sreported-- GB2,048llama.cpp
Radeon AI Pro R9700 32GB4-bit25.4 tok/sreported555- GB2,048llama.cpp
RTX 3060 12GB ×24-bit15.2 tok/smeasured10,91121.5 GB98,304llama.cpp
RTX 3090 24GB4-bit38.8 tok/sreported-- GB2,048llama.cpp
RTX 3090 24GB ×26-bit14.4 tok/sreported2,26320.7 GB32,768llama.cpp
RTX 5090 32GB6-bit51.7 tok/sreported155- GB2,048llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.