community pool / model answer

DeepSeek-V4-Flash

chat · deepseek-ai/DeepSeek-V4-Flash

28.5 tok/s · model median · 36 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
GB10 Grace Blackwell 128GB8-bit45.7 tok/sreported152209 GB200,000vllm
GB10 Grace Blackwell 256GB4-bit26.9 tok/sreported495- GB393,216vllm
GB10 Grace Blackwell 256GB8-bit36 tok/sreported1,20299.2 GB500,000vllm
M3 Ultra 512GB2-bit33 tok/sreported300- GB16,384mlx
M5 Max 128GB2-bit19 tok/smeasured70,150- GB128,000llama.cpp
Radeon 8060S 103.1GB2-bit19 tok/sreported65983 GB4,096hipfire
Radeon AI Pro R9700 32GB ×31-bit9.2 tok/smeasured4,40381.6 GB8,192llama.cpp
Radeon AI Pro R9700 32GB ×42-bit53.7 tok/sreported5,602- GB2,052hipfire
RTX 3090 24GB ×42-bit35.3 tok/sreported1,28383.9 GB2,048llama.cpp
RTX 5090 32GB ×42-bit212 tok/sreported55124 GB262,144vllm
RTX PRO 6000 Blackwell 96GB3-bit80.8 tok/sreported9,24094.5 GB262,144llama.cpp
RTX PRO 6000 Blackwell 96GB ×24-bit160.9 tok/sreported1,130188.8 GB500,000vllm
RTX PRO 6000 Blackwell 96GB ×28-bit261.8 tok/sreported4,830184 GB262,000vllm
Ryzen AI Max 395 103GB2-bit34.9 tok/sreported1,159- GB4,096hipfire
Ryzen AI Max 395 128GB2-bit15.2 tok/smeasured-- GB65,536dwarfstar, hipfire

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.