community pool / model answer

Llama 3.1 8B Instruct GGUF

chat · unsloth/Llama-3.1-8B-Instruct-GGUF

No data yet · 6 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
L4 24GB (modal)4-bit44.8 tok/smeasured6175.4 GB4,096llama.cpp
A10 24GB (modal)4-bit75.8 tok/smeasured3865.4 GB4,096llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.