community pool / model answer

DeepSeek-V4-Flash-0731

chat · deepseek-ai/DeepSeek-V4-Flash-0731

48.7 tok/s · model median · 41 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
GB10 Grace Blackwell 256GB8-bit24.9 tok/smeasured185- GB2,048vllm
Radeon AI Pro R9700 32GB2-bit60.6 tok/sreported2,00925.2 GB8,192lucebox
Radeon AI Pro R9700 32GB4-bit50.5 tok/sreported6,06328.1 GB32,768llama.cpp
Radeon AI Pro R9700 32GB ×31-bit37.3 tok/smeasured1,880- GB32,768llama.cpp
RTX 3090 24GB ×88-bit93.3 tok/sreported191184.4 GB327,680vllm
RTX PRO 6000 Blackwell 96GB ×28-bit257.4 tok/sreported8394.5 GB262,144vllm
RTX PRO 6000 Blackwell 96GB ×44-bit231.5 tok/sreported50- GB2,048vllm
RTX PRO 6000 Blackwell 96GB ×48-bit296.7 tok/smeasured90- GB2,048vllm
Ryzen AI Max 395 128GB1-bit21.8 tok/sreported3,305- GB782llama.cpp
Ryzen AI Max 395 128GB2-bit28.3 tok/sreported15,107- GB782lucebox

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.