community pool / model answer

Qwen3.6-35B-A3B-NVFP4

chat · nvidia/Qwen3.6-35B-A3B-NVFP4

233.4 tok/s · model median · 19 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
GB10 Grace Blackwell 128GB4-bit106 tok/smeasured135- GB262,144vllm
RTX 3090 24GB ×24-bit233.4 tok/smeasured4541.4 GB262,144vllm
RTX 3090 24GB ×34-bit231.3 tok/sreported-- GB2,048vllm
RTX PRO 6000 Blackwell 96GB4-bit330.3 tok/smeasured5,579- GB262,144vllm
RTX PRO 6000 Blackwell 96GB ×44-bit325.9 tok/sreported39- GB2,048vllm

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.