community pool / model answer

Qwopus3.5-27B-v3

chat · Jackrong/Qwopus3.5-27B-v3

12.9 tok/s · model median · 23 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
RTX 3090 24GB4-bit25.8 tok/smeasured7,9510 GB131,072llama.cpp
Ryzen AI Max 395 128GB4-bit12.9 tok/smeasured-17.5 GB32,768llama.cpp
Ryzen AI Max 395 128GB5-bit11.3 tok/smeasured-19.9 GB32,768llama.cpp
Ryzen AI Max 395 128GB6-bit9.9 tok/sreported-23.5 GB32,768llama.cpp
Ryzen AI Max 395 128GB8-bit7.9 tok/sreported-29.2 GB32,768llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.