community pool / model answer

Gemma-4-31B-it-abliterated-GGUF

chat · LiconStudio/Gemma-4-31B-it-abliterated-GGUF

87.7 tok/s · model median · 1 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
RTX 4090 24GB4-bit87.7 tok/sreported-23.3 GB70,080llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.