community-verified · measured pool data
Can the Ryzen AI Max 395 128GB run Llama-3.1-Nemotron-70B-Instruct-HF?
Yes — community-measured at 5.2 tok/s (2 runs at 4-bit).
Verified single-stream runs on the Ryzen AI Max 395 128GB from the localmaxxing public pool. Every number on this page is measured — nothing here is estimated.
| Quantization | Median tok/s | TTFT (ms) | Peak VRAM | Max context tested | Engines |
|---|---|---|---|---|---|
| 4-bit | 5.2 | — | — | 16,384 | llama.cpp |