community pool / model answer

Qwen3.5-2B

chat · Qwen/Qwen3.5-2B

105.9 tok/s · model median · 4 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
AMD Ryzen 7 8745H w/ Radeon 780M4-bit45.9 tok/sreported4262.5 GB8,192llama.cpp
AMD Ryzen 7 8745H w/ Radeon 780M Graphics4-bit45.9 tok/sreported424- GB65,536llama.cpp
RTX 3060 12GB4-bit166.2 tok/sreported7063.5 GB4,608llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.