community pool / model answer

Ornith-1.5-9B-GGUF

chat · ornith-ai/Ornith-1.5-9B-GGUF

73.9 tok/s · model median · 4 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
Arc Pro B70 32GB8-bit49.6 tok/sreported-- GB8,192llama.cpp
RTX 3060 12GB4-bit65 tok/sreported60- GB65,536llama.cpp
RTX 3080 12GB5-bit92 tok/sreported7629.5 GB262,144llama.cpp
RTX 3080 12GB6-bit82.8 tok/sreported8008.5 GB131,072llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.