community-verified · measured pool data
Can the RTX 5080 16GB run Meta-Llama-3-8B-Instruct?
Yes — community-measured at 129.5 tok/s (6 runs at 4-bit).
Verified single-stream runs on the RTX 5080 16GB from the localmaxxing public pool. Every number on this page is measured — nothing here is estimated.
| Quantization | Median tok/s | TTFT (ms) | Peak VRAM | Max context tested | Engines |
|---|---|---|---|---|---|
| 4-bit | 129.5 | 79.5 | — | 4,096 | llama.cpp |