community pool / model answer
DeepSeek-V4.1-Flash
chat · deepseek-ai/DeepSeek-V4.1-Flash
72 tok/s · model median · 13 runs in the model record.
| Reference rig | Quantization | Median result | Basis | TTFT (ms) | Peak VRAM | Max context | Engines |
|---|---|---|---|---|---|---|---|
| GB10 Grace Blackwell 512GB | 4-bit | 59.1 tok/s | measured | 391 | - GB | 430,080 | vllm |
| GB10 Grace Blackwell 512GB | 8-bit | 31.7 tok/s | reported | 3,328 | - GB | 2,048 | vllm |
| RTX PRO 6000 Blackwell Max-Q Workstation Edition 96GB ×4 | 8-bit | 303.7 tok/s | measured | 221 | - GB | 524,288 | vllm |
Context tested
The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.