community-verified · measured pool data

Can the RTX 5090 32GB run Qwen3.8-27B-GGUF?

Yes — community-measured at 66.2 tok/s (4 runs at 4-bit).

Verified single-stream runs on the RTX 5090 32GB from the localmaxxing public pool. Every number on this page is measured — nothing here is estimated.

Quantization Median tok/s TTFT (ms) Peak VRAM Max context tested Engines
3-bit 79.4 381.6 17.3 GB 131,072 llama.cpp
4-bit 66.2 397.6 21.3 GB 131,072 llama.cpp, lmstudio
5-bit 60.2 424.4 23.2 GB 131,072 llama.cpp
6-bit 51.9 423.3 26.5 GB 65,536 llama.cpp, lmstudio
Go deeper
I have a goal — find models that fit my use case → I have hardware — browse everything my rig can run → All machine × model answer pages →