community pool / model answer

Llama-3.2-1B-Instruct

chat · meta-llama/Llama-3.2-1B-Instruct

184 tok/s · model median · 7 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
Arc Pro B60 24GB8-bit189.4 tok/sreported64- GB4,096llama.cpp
Arc Pro B60 24GB ×28-bit189.3 tok/sreported64- GB4,096llama.cpp
AMD Ryzen 7 7840HS2-bit91.1 tok/sreported-- GB2,048llama.cpp
AMD Ryzen 7 7840HS4-bit68.2 tok/sreported-- GB2,048llama.cpp
RTX 5080 16GB8-bit448.1 tok/sreported14- GB4,096llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.