community pool / model answer

Qwen-AgentWorld-35B-A3B-GGUF

chat · unsloth/Qwen-AgentWorld-35B-A3B-GGUF

69.8 tok/s · model median · 19 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
RTX 3060 12GB ×24-bit78.3 tok/smeasured-21.6 GB32,768llama.cpp
Ryzen AI Max 395 128GB4-bit34.9 tok/sreported37,26721.8 GB65,536llama.cpp
Ryzen AI Max 395 64GB4-bit60.5 tok/smeasured1,57621.5 GB8,192llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.