community pool / model answer

Ornith-1.0-35B

chat · ornith-ai/Ornith-1.0-35B

71.8 tok/s · model median · 12 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
Arc Pro B70 32GB4-bit74.3 tok/sreported73- GB262,144llama.cpp
Arc Pro B70 32GB5-bit69.3 tok/smeasured-- GB262,144llama.cpp
M3 Ultra 512GB4-bit89.1 tok/sreported111- GB8,192llama.cpp
Radeon AI Pro R9700 32GB ×38-bit74 tok/sreported36642.9 GB787llama.cpp
Ryzen AI Max 395 128GB8-bit54.1 tok/sreported1,244- GB787llama.cpp
Ryzen AI Max 395 64GB4-bit60.6 tok/smeasured97021.4 GB8,192llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.