community pool / model answer

DeepSeek-V4-Flash-0731-GGUF

chat · unsloth/DeepSeek-V4-Flash-0731-GGUF

33.9 tok/s · model median · 7 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
RTX 3090 24GB ×42-bit41 tok/sreported18,77490.3 GB131,072llama.cpp
RTX 3090 24GB ×74-bit42.4 tok/sreported916146.5 GB262,144llama.cpp
RTX 5090 32GB1-bit17.8 tok/sreported250- GB2,048llama.cpp
RTX 5090 32GB8-bit10.4 tok/sreported271- GB2,048llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.