community-verified · measured pool data
Can the RTX 3060 12GB run Llama-3.2-3B-Instruct?
Yes — community-measured at 95.8 tok/s (39 runs at 4-bit).
Verified single-stream runs on the RTX 3060 12GB from the localmaxxing public pool. Every number on this page is measured — nothing here is estimated.
| Quantization | Median tok/s | TTFT (ms) | Peak VRAM | Max context tested | Engines |
|---|---|---|---|---|---|
| 4-bit | 95.8 | 681.3 | — | 131,072 | llama.cpp |