community pool / model answer

Muse-Glimmer-30B-GGUF

chat · meta-models/Muse-Glimmer-30B-GGUF

28.2 tok/s · model median · 15 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
Arc Pro B70 32GB4-bit28.8 tok/smeasured21022 GB2,048llama.cpp
Radeon AI Pro R9700 32GB4-bit82.5 tok/sreported25923.7 GB32,768llama.cpp
RTX 3080 12GB4-bit6 tok/sreported3,8408.8 GB16,384llama.cpp
RTX PRO 6000 Blackwell 96GB4-bit167.5 tok/sreported56124.3 GB131,072llama.cpp
Ryzen AI Max 395 64GB4-bit12.9 tok/smeasured4,55816.4 GB8,192llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.