community pool / model answer

Qwen3.5-35B-A3B

chat · Qwen/Qwen3.5-35B-A3B

85.6 tok/s · model median · 7 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
Arc Pro B70 32GB4-bit62.8 tok/sreported-- GB262,144llama.cpp
Arc Pro B70 32GB5-bit61.5 tok/sreported-- GB262,144llama.cpp
GTX 1080 Ti 11GB4-bit9.3 tok/sreported-9.4 GB8,192llama.cpp
M5 Max 128GB4-bit122 tok/sreported-- GB262,144ollama
RX 7900 XTX 24GB4-bit140.6 tok/smeasured7522.1 GB4,096hipfire

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.