community pool / model answer

Gemopus-4-26B-A4B-it

chat · Jackrong/Gemopus-4-26B-A4B-it

55 tok/s · model median · 4 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
Ryzen AI Max 395 128GB4-bit64.2 tok/sreported-23.7 GB128,000llama.cpp
Ryzen AI Max 395 128GB8-bit45.8 tok/sreported-33.1 GB128,000llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.