community pool / model answer

gpt-oss-20b-GGUF

chat · ggml-org/gpt-oss-20b-GGUF

65.9 tok/s · model median · 2 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
M1 Max 64GB4-bit74.6 tok/sreported559- GB2,048llama.cpp
RTX 3080 12GB4-bit57.1 tok/sreported1,5589.3 GB32,768llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.