community-verified · measured pool data

Can the RTX3060+RTX2080-pooled-20GB 20GB run Llama-3.1-8B-Instruct?

Yes — community-measured at 105 tok/s (2 runs at 4-bit).

Verified single-stream runs on the RTX3060+RTX2080-pooled-20GB 20GB from the localmaxxing public pool. Every number on this page is measured — nothing here is estimated.

Quantization Median tok/s TTFT (ms) Peak VRAM Max context tested Engines
4-bit 105 5 GB 2,048 ollama
Go deeper
I have a goal — find models that fit my use case → I have hardware — browse everything my rig can run → All machine × model answer pages →