community pool / model answer

gemma-4-E4B-it-GGUF

chat · unsloth/gemma-4-E4B-it-GGUF

68.8 tok/s · model median · 35 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
Minisforum UM790 Pro 64GB4-bit21.7 tok/sreported-- GB2,048llama.cpp
Radeon 890M 32GB4-bit18.7 tok/sreported1,350- GB2,048llama.cpp
RTX 2080 Ti 11GB4-bit106 tok/sreported-- GB2,048llama.cpp
RTX 2080 Ti 11GB6-bit89 tok/sreported-- GB2,048llama.cpp
RTX 2080 Ti 11GB8-bit80 tok/sreported-- GB2,048llama.cpp
RTX 3050 8GB5-bit46.4 tok/sreported169- GB128,000llama.cpp
RTX 3060 12GB4-bit69.9 tok/smeasured1,0663.5 GB131,072llama.cpp
RTX 3060 12GB ×24-bit69.3 tok/smeasured5,1855.3 GB131,072llama.cpp
Ryzen AI Max 385 24GB4-bit18.7 tok/sreported1,350- GB2,048llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.