community pool / model answer

Gemma-4-26B-A4B-NVFP4

chat · nvidia/Gemma-4-26B-A4B-NVFP4

482.3 tok/s · model median · 4 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
GB10 Grace Blackwell 128GB4-bit227.5 tok/sreported200- GB2,048vllm
M1 Max 64GB4-bit52.7 tok/sreported-- GB2,048ollama
RTX PRO 6000 Blackwell Server Edition 96GB4-bit128.3 tok/sreported-87.5 GB65,536vllm
RTX PRO 6000 Blackwell Workstation Edition 96GB4-bit36.9 tok/sreported60- GB131,072vllm

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.