community pool / model answer

Ornith-1.0-9B

chat · ornith-ai/Ornith-1.0-9B

25.7 tok/s · model median · 44 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
Arc Pro B70 32GB6-bit50.3 tok/sreported58- GB262,144llama.cpp
M3 Ultra 512GB4-bit77.5 tok/sreported104- GB8,192llama.cpp
Radeon AI Pro R9700 32GB ×316-bit25.7 tok/smeasured39425.8 GB787llama.cpp
RX 570 4GB4-bit5.3 tok/smeasured411- GB2,048llama.cpp
RX 570 4GB ×24-bit23.8 tok/smeasured157- GB2,048llama.cpp
RX 6700 XT 12GB8-bit37.2 tok/smeasured-- GB4,736llama.cpp
RX 7700 XT 12GB5-bit30.6 tok/sreported87- GB2,048vllm

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.