community pool / model answer

diffusiongemma-26B-A4B-it-FP8-dynamic

chat · RedHatAI/diffusiongemma-26B-A4B-it-FP8-dynamic

773.4 tok/s · model median · 2 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
H100 80GB8-bit1,369.3 tok/sreported260- GB8,192vllm
RTX 3090 24GB ×28-bit177.6 tok/sreported1,675- GB262,144vllm

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.