community pool / model answer

gemma-4-12B-it-qat-q4_0-gguf

chat · google/gemma-4-12B-it-qat-q4_0-gguf

18.9 tok/s · model median · 8 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
RTX 3080 12GB4-bit71.2 tok/sreported9688.6 GB262,144llama.cpp
RTX 5090 32GB4-bit130.8 tok/sreported100- GB2,048lmstudio
RX 570 4GB ×24-bit16.7 tok/smeasured737- GB2,048llama.cpp
RX 7700 XT 12GB4-bit40.4 tok/sreported226- GB2,048vllm

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.