community pool / model answer

Bonsai-27B-gguf

chat · prism-ml/Bonsai-27B-gguf

31.5 tok/s · model median · 36 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
RTX 3060 12GB ×24-bit29.3 tok/smeasured121.9 GB262,144llama.cpp
RTX 3080 12GB1-bit72.9 tok/sreported2,0545.1 GB32,768llama.cpp
RTX 3090 24GB1-bit39.8 tok/smeasured31,9045.6 GB262,144llama.cpp
RTX 5090 32GB1-bit136.7 tok/sreported82- GB2,048lmstudio
RX 7700 XT 12GB1-bit44 tok/sreported117- GB2,048vllm
Ryzen AI Max 395 128GB1-bit33.6 tok/smeasured26,241- GB262,144llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.