community-verified · measured pool data
Can the Multi-GPU 72GB ×3 run Ling-3.0-flash?
Yes — community-measured at 53.4 tok/s (5 runs at 4-bit).
Verified single-stream runs on the Multi-GPU 72GB ×3 from the localmaxxing public pool. Every number on this page is measured — nothing here is estimated.
| Quantization | Median tok/s | TTFT (ms) | Peak VRAM | Max context tested | Engines |
|---|---|---|---|---|---|
| 4-bit | 53.4 | 653.6 | 68.4 GB | 262,144 | llama.cpp |