community pool / model answer

Qwen3-32B

chat · Qwen/Qwen3-32B

22.8 tok/s · model median · 4 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
M1 Max 64GB4-bit9 tok/sreported114- GB2,048ollama
Radeon AI Pro R9700 32GB ×28-bit22.9 tok/sreported71- GB32,768vllm
Radeon AI Pro R9700 32GB ×38-bit22.8 tok/sreported6034.2 GB786vllm
Radeon AI Pro R9700 57GB4-bit3.2 tok/sreported8,835- GB2,048llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.