community pool / model answer

Agents-A1-Q4_K_M-GGUF

chat · InternScience/Agents-A1-Q4_K_M-GGUF

74.6 tok/s · model median · 17 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
RTX 3060 12GB ×24-bit91.7 tok/smeasured3,12620.9 GB65,536llama.cpp
RX 570 4GB4-bit8.7 tok/sreported644- GB2,048llama.cpp
RX 570 4GB ×24-bit16.4 tok/sreported174- GB2,048llama.cpp
Ryzen AI Max 395 128GB4-bit62 tok/sreported46,678- GB262,144llama.cpp
Ryzen AI Max 395 64GB4-bit74.6 tok/smeasured1,59220.4 GB8,192llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.