community pool / model answer

Laguna-S-2.1

chat · poolside/Laguna-S-2.1

41.9 tok/s · model median · 47 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
Multi-GPU 72GB ×34-bit60.3 tok/smeasured32765.1 GB278,528llama.cpp
Radeon AI Pro R9700 32GB ×34-bit48.3 tok/smeasured78065.5 GB131,072llama.cpp
RTX PRO 6000 Blackwell 96GB ×416-bit126.6 tok/sreported50- GB2,048vllm
Ryzen AI Max 395 128GB4-bit40.7 tok/smeasured2,811- GB262,144llama.cpp
Ryzen AI Max 395 64GB2-bit43.7 tok/smeasured2,64035.9 GB10,240llama.cpp
Ryzen AI Max 395 64GB3-bit32.8 tok/smeasured2,82352 GB10,240llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.