community pool / model answer

gemma-3-12b-it

chat · google/gemma-3-12b-it

44 tok/s · model median · 41 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
RTX 2080 8GB4-bit21 tok/sreported-8.1 GB2,048ollama
RTX 3060 12GB4-bit40 tok/sreported-8.1 GB2,048ollama
RTX 3060 12GB ×24-bit34.3 tok/smeasured8,28912.2 GB131,072llama.cpp
RTX 3080 12GB4-bit51.5 tok/smeasured-8.6 GB2,048ollama
RTX 3090 24GB4-bit59.6 tok/smeasured3,320- GB131,072llama.cpp
RTX3060+RTX2080-pooled-20GB 20GB4-bit44 tok/smeasured-8.1 GB2,048ollama
Ryzen AI Max 395 64GB4-bit24.6 tok/sreported1,7719.1 GB8,192llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.