community pool / model answer

Laguna-XS-2.1-GGUF

chat · poolside/Laguna-XS-2.1-GGUF

103.9 tok/s · model median · 25 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
M1 Max 64GB4-bit62.3 tok/sreported705- GB2,048llama.cpp
RTX 3090 24GB3-bit110.7 tok/smeasured-19.8 GB262,144llama.cpp
RTX 3090 24GB4-bit125.2 tok/smeasured-21.5 GB131,072llama.cpp
RTX 4070 12GB4-bit55.1 tok/sreported72,847- GB2,048llama.cpp
RX 570 4GB ×24-bit16.8 tok/sreported206- GB2,048llama.cpp
Ryzen AI Max 395 64GB4-bit92.4 tok/sreported96119.8 GB8,192llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.