community pool / model answer

Gemma-4-31B-IT-NVFP4

chat · nvidia/Gemma-4-31B-IT-NVFP4

31.7 tok/s · model median · 1 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
RTX PRO 6000 Blackwell 72GB4-bit31.7 tok/sreported4167.8 GB131,072vllm

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.