community pool / model answer

NVIDIA-Nemotron-Nano-9B-v2

chat · nvidia/NVIDIA-Nemotron-Nano-9B-v2

64.3 tok/s · model median · 22 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
RTX 3060 12GB ×24-bit45.2 tok/smeasured6,1448.6 GB131,072llama.cpp
RTX 3060 + RTX 2080 20GB4-bit50 tok/sreported-5.2 GB2,048llama.cpp
RTX 3060 Ti 8GB4-bit64.3 tok/sreported292- GB2,048llama.cpp
RTX 3080 12GB4-bit98 tok/sreported-5.2 GB2,048llama.cpp
RTX 3090 24GB4-bit73.7 tok/smeasured3,388- GB131,072llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.