community pool / model answer

glm-4-9b-chat-abliterated

chat · byroneverson/glm-4-9b-chat-abliterated

35.4 tok/s · model median · 5 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
Radeon AI Pro R9700 32GB ×34-bit65.8 tok/sreported1799.1 GB783llama.cpp
RX 6700 XT 12GB8-bit35 tok/smeasured-- GB4,736llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.