community pool / model answer

gemma-4-26B-A4B-it

chat · google/gemma-4-26B-A4B-it

60.6 tok/s · model median · 43 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
Arc Pro B70 32GB8-bit43.3 tok/sreported393- GB2,048llama.cpp
AMD Ryzen 7 8745H w/ Radeon 780M4-bit30.6 tok/sreported1,5344 GB256,000llama.cpp
M1 Max 64GB8-bit49 tok/sreported462- GB2,048llama.cpp, ollama
M3 Ultra 512GB4-bit76.9 tok/sreported279- GB2,048mlx
Radeon AI Pro R9700 32GB ×24-bit87.3 tok/sreported375- GB794llama.cpp
Radeon AI Pro R9700 32GB ×28-bit99.3 tok/smeasured-- GB262,144llama.cpp
Radeon AI Pro R9700 32GB ×38-bit73.7 tok/smeasured29835.4 GB794llama.cpp
RTX 3090 24GB ×28-bit154.1 tok/sreported4245.3 GB262,144vllm
RTX 4070 12GB4-bit20.4 tok/smeasured1,661- GB2,048vllm
RTX 5090 32GB8-bit282 tok/smeasured30- GB2,048vllm
RTX PRO 6000 Blackwell 96GB8-bit276.7 tok/sreported28- GB8,192vllm
RX 570 4GB ×24-bit12.5 tok/smeasured760- GB2,048llama.cpp
Ryzen AI Max 395 128GB4-bit60.7 tok/smeasured-21.5 GB32,768llama.cpp
Ryzen AI Max 395 128GB6-bit55.5 tok/sreported-23.8 GB16,384llama.cpp
Ryzen AI Max 395 128GB8-bit46.1 tok/sreported74627.8 GB16,384llama.cpp
Tesla P100-PCIE-16GB 16GB ×25-bit54.8 tok/sreported-- GB2,048llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.