community pool / model answer

Qwen3.6-35B-A3B

chat · unsloth/Qwen3.6-35B-A3B

57.6 tok/s · model median · 16 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
M3 Ultra 512GB3-bit84.6 tok/sreported224- GB8,192lmstudio
Radeon RX 7900 XTX 24GB3-bit104.5 tok/sreported-15.7 GB640llama.cpp
RTX 3090 24GB4-bit91.8 tok/sreported11521.9 GB32,768llama.cpp
RTX 4060 Ti 16GB3-bit60.2 tok/sreported4,000- GB262,144llama.cpp
RTX 4070 12GB4-bit57.4 tok/smeasured6811.6 GB131,072llama.cpp
Ryzen AI Max 395 128GB4-bit57.5 tok/smeasured-27.9 GB131,072llama.cpp
Ryzen AI Max 395 128GB8-bit52.8 tok/smeasured-40.4 GB128,000llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.