community pool / model answer

Qwen3-Coder-30B-A3B-Instruct

code · Qwen/Qwen3-Coder-30B-A3B-Instruct

80.4 tok/s · model median · 21 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
Arc Pro B70 32GB ×24-bit93.3 tok/sreported226- GB2,048llama.cpp
M1 Max 64GB4-bit72.3 tok/sreported24- GB2,048ollama
M3 Ultra 512GB4-bit100 tok/sreported79- GB8,192llama.cpp
Radeon AI Pro R9700 32GB ×38-bit79 tok/sreported204- GB686llama.cpp
RX 7900 XTX 24GB4-bit43.8 tok/sreported2,11624.2 GB2,048llama.cpp
Ryzen AI Max 395 128GB5-bit80.4 tok/smeasured-23.9 GB100,000llama.cpp
Ryzen AI Max 395 64GB4-bit31.9 tok/smeasured69930.3 GB131,072llama.cpp
Tesla V100 32GB ×26-bit100.5 tok/sreported-58 GB65,536llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.