community pool / model answer

GLM-5.3-Flash-GGUF

chat · unsloth/GLM-5.3-Flash-GGUF

22.6 tok/s · model median · 4 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
Radeon AI Pro R9700 32GB1-bit17.4 tok/sreported73292.2 GB32,768llama.cpp
Radeon AI Pro R9700 32GB2-bit16.6 tok/sreported786101.9 GB32,768llama.cpp
RTX 3090 24GB ×84-bit35 tok/sreported2,167166.8 GB204,800llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.