community pool / model answer

Qwen3.5-0.8B

chat · Qwen/Qwen3.5-0.8B

445.7 tok/s · model median · 14 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
AMD Ryzen 7 7840HS4-bit77.8 tok/sreported-- GB2,048llama.cpp
Radeon AI Pro R9700 32GB4-bit562.2 tok/smeasured44- GB128hipfire
RTX 3060 12GB4-bit209.8 tok/sreported59- GB4,096llama.cpp
RTX 3060 12GB8-bit242.8 tok/sreported4333 GB4,608llama.cpp
RTX 5060 Ti 16GB4-bit280.9 tok/sreported48- GB4,096llama.cpp
RTX PRO 6000 Blackwell 96GB16-bit778.4 tok/sreported37- GB2,048vllm
RX 5700 XT 8GB4-bit273.6 tok/sreported347- GB128hipfire
RX 6900 XT 16GB4-bit346.1 tok/sreported81- GB128hipfire
RX 7900 XTX 24GB4-bit550.2 tok/sreported69- GB128hipfire
RX 9070 XT 16GB4-bit545.3 tok/sreported67- GB128hipfire
Ryzen AI Max 395 103GB4-bit291.4 tok/sreported69- GB128hipfire

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.