community pool / model answer

Step-3.7-Flash-GGUF

chat · stepfun-ai/Step-3.7-Flash-GGUF

30.7 tok/s · model median · 9 runs in the model record.

Reference rigQuantizationMedian resultBasisTTFT (ms)Peak VRAMMax contextEngines
Arc Pro B70 32GB ×44-bit20.8 tok/sreported-95.6 GB32,768llama.cpp
GB10 Grace Blackwell 128GB4-bit23 tok/sreported8,470- GB2,048llama.cpp
RTX 3090 24GB ×84-bit52.7 tok/sreported199- GB262,144llama.cpp
Ryzen AI Max 395 128GB4-bit22 tok/sreported36,084- GB65,536llama.cpp
Ryzen AI Max 395 64GB1-bit30.8 tok/smeasured3,33453.8 GB4,096llama.cpp

Context tested

The table above is the tested context for each exact rig and quantization cell. Empty fields remain empty rather than being filled with an estimate.