community pool / model answer

Qwen3.5-4B

chat · Qwen/Qwen3.5-4B

61 tok/s · model median · 20 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
AMD Ryzen 7 8745H w/ Radeon 780M4-bit28 tok/smeasured1,5373.5 GB65,536llama.cpp
AMD Ryzen 7 8745H w/ Radeon 780M Graphics4-bit22 tok/sreported961- GB65,536llama.cpp
Intel(R) Core(TM) Ultra 7 155H4-bit14.6 tok/sreported-- GB2,048llama.cpp
GTX 1650 4GB4-bit30.6 tok/sreported14,583- GB4,096llama.cpp
RTX 3060 12GB4-bit76.5 tok/smeasured1,124- GB131,072llama.cpp
RX 6700 10GB4-bit76.6 tok/sreported2,1414 GB4,096llama.cpp
Ryzen AI Max 395 128GB4-bit61.3 tok/smeasured-4.3 GB32,768llama.cpp
Ultra 7 155H 32GB4-bit12 tok/sreported-- GB2,048llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.