community pool / model answer

Qwen3.5-9B-GGUF

chat · unsloth/Qwen3.5-9B-GGUF

51.9 tok/s · model median · 15 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
RTX 3060 12GB4-bit52.2 tok/sreported2,4926.5 GB32,768llama.cpp
RTX 3060 12GB ×24-bit51 tok/smeasured5,6407.3 GB131,072llama.cpp
RTX 3090 24GB ×24-bit110.1 tok/sreported-- GB8,192llama.cpp
RTX 5070 Ti 16GB4-bit124.1 tok/sreported-- GB8,192llama.cpp
Tesla P100-PCIE-16GB 16GB5-bit32.4 tok/sreported-- GB2,048llama.cpp
Tesla P100-PCIE-16GB 16GB ×25-bit49.9 tok/sreported-- GB2,048llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.