community-verified · measured pool data
Can the RTX A5000 24GB run Qwen2.5-1.5B-Instruct-GGUF?
Yes — community-measured at 320.6 tok/s (2 runs at 4-bit).
Verified single-stream runs on the RTX A5000 24GB from the localmaxxing public pool. Every number on this page is measured — nothing here is estimated.
| Quantization | Median tok/s | TTFT (ms) | Peak VRAM | Max context tested | Engines |
|---|---|---|---|---|---|
| 4-bit | 320.6 | — | 1.6 GB | 4,096 | llama.cpp |
| 8-bit | 268 | — | 2.2 GB | 4,096 | llama.cpp |