community pool / model answer

Muse-Glimmer-30B

chat · meta-models/Muse-Glimmer-30B

41.9 tok/s · model median · 17 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
Arc Pro B70 32GB4-bit26.8 tok/smeasured-- GB131,072llama.cpp
Radeon AI Pro R9700 32GB ×34-bit52.8 tok/sreported1,372- GB1,923llama.cpp
Radeon AI Pro R9700 32GB ×38-bit21.3 tok/smeasured875- GB1,923llama.cpp
RTX PRO 6000 Blackwell 96GB16-bit48.6 tok/sreported23952.3 GB131,072llama.cpp, vllm
Ryzen AI Max 395 64GB4-bit12.9 tok/sreported4,68916.4 GB8,192llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.