community pool / model answer

Qwen3.8-27B-QUASAR-NVFP4

chat · QUASAR-QAT/Qwen3.8-27B-QUASAR-NVFP4

263.7 tok/s · model median · 2 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
RTX 5070 Ti 16GB ×24-bit100.1 tok/sreported-29 GB2,048sglang
Tesla V100 32GB ×44-bit427.3 tok/sreported315117.4 GB262,144vllm

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.