community pool / model answer

Qwen3.5-27B

chat · Qwen/Qwen3.5-27B

25.1 tok/s · model median · 18 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
Arc Pro B70 32GB4-bit18 tok/sreported48620 GB8,192llama.cpp
Arc Pro B70 32GB5-bit25.1 tok/sreported-- GB204,800llama.cpp
GB10 Grace Blackwell 128GB4-bit11.5 tok/smeasured658113.5 GB20,000vllm
GTX 1080 Ti 11GB4-bit2.5 tok/sreported-10.3 GB8,192llama.cpp
Radeon 8060S Graphics 103GB4-bit104.5 tok/sreported33818 GB4,096hipfire
Radeon 8060S Graphics (Strix Halo APU, gfx1151) 96GB4-bit14.8 tok/sreported48016 GB65,536hipfire
Radeon AI Pro R9700 32GB4-bit196.2 tok/sreported13317.9 GB4,096hipfire
Radeon AI Pro R9700 34GB4-bit286.6 tok/sreported11117.5 GB4,096hipfire
RX 7900 XTX 24GB4-bit191.5 tok/smeasured13617 GB4,096hipfire

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.