community pool / model answer

DeepSeek-R1-Distill-Qwen-14B

chat · deepseek-ai/DeepSeek-R1-Distill-Qwen-14B

38 tok/s · model median · 10 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
RTX 3090 24GB4-bit41 tok/smeasured4,8000 GB131,072llama.cpp
RTX 5000 16GB4-bit24.4 tok/sreported-15.2 GB57,344llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.