community pool / model answer

Llama-3.2-3B-Instruct

chat · meta-llama/Llama-3.2-3B-Instruct

75.9 tok/s · model median · 44 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
M1 Max 64GB16-bit50.9 tok/sreported-- GB640llama.cpp
RTX 3060 12GB4-bit95.8 tok/smeasured681- GB131,072llama.cpp
RX 6700 XT 12GB16-bit48.3 tok/smeasured-- GB4,736llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.