community pool / model answer

Qwen3.6-27B-int4-AutoRound

chat · Lorbus/Qwen3.6-27B-int4-AutoRound

66.4 tok/s · model median · 21 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
Arc Pro B70 32GB4-bit41.3 tok/smeasured-- GB4,096vllm
Arc Pro B70 32GB ×24-bit48.3 tok/smeasured-- GB4,096vllm
RTX 3090 24GB4-bit74.5 tok/smeasured48022.5 GB32,768vllm
RTX 3090 24GB ×24-bit78.2 tok/smeasured7922.8 GB262,144vllm
RTX 4090 24GB16-bit52.6 tok/sreported88- GB2,048vllm
RTX 5060 Ti 16GB ×24-bit69.3 tok/sreported94- GB2,048vllm

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.