community pool / model answer

gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF

chat · yuxinlu1/gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF

32.3 tok/s · model median · 6 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
RTX 4070 12GB4-bit39.7 tok/smeasured570- GB2,048vllm
Ryzen AI Max 395 128GB4-bit24.7 tok/sreported2189.4 GB4,096llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.