community pool / model answer

gemma-4-26B-A4B-it-qat-GGUF

chat · unsloth/gemma-4-26B-A4B-it-qat-GGUF

95.3 tok/s · model median · 42 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
M1 Max 64GB4-bit76.4 tok/sreported50- GB2,048ollama
M5 Pro 24GB4-bit121.6 tok/sreported147- GB2,048llama.cpp
RTX 3050 8GB4-bit38 tok/sreported952- GB28,000llama.cpp
RTX 3060 12GB ×24-bit87.8 tok/smeasured2,37916 GB131,072llama.cpp
RTX 3090 24GB4-bit138.3 tok/smeasured3,4330 GB131,072llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.