community-verified · measured pool data

Can the RTX 3090 24GB run Qwen3.6-27B-MTP-GGUF?

Yes — community-measured at 49.9 tok/s (16 runs at 4-bit).

Verified single-stream runs on the RTX 3090 24GB from the localmaxxing public pool. Every number on this page is measured — nothing here is estimated.

Quantization Median tok/s TTFT (ms) Peak VRAM Max context tested Engines
1-bit 49.2 2.5 GB 196,608 llama.cpp
4-bit 49.9 465.2 20.2 GB 131,072 llama.cpp
Go deeper
I have a goal — find models that fit my use case → I have hardware — browse everything my rig can run → All machine × model answer pages →