Skip to content
Bencher

Models · StepFun · Out sinceReleased 28 May 2026

Step 3.7 Flash

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableCan run on your own serversOpen weights

Step 3.7 Flash is made by StepFun. We don't have enough test results yet to rank it. It's cheap to use. Its makers have published it, so you can run it on your own servers.

198B sparse MoE VLM (196B LM + 1.8B vision encoder), ~11B active, 256K context, Apache 2.0. Price from the model card (StepFun platform: $0.20 in cache-miss / $0.04 cache-hit / $1.15 out). HF repo created 2026-05-23; OpenRouter listed 2026-05-28.

Writing & creativity
—
Research & analysis
—
Coding
—
Price
Cheap$0.20 / $1.15
Price per 1M (blended)Blended / 1M
$0.44
MemoryContext
262K
Longest answerMax output
—
Test resultsResults
0 (0 independent0 indep.)
Out sinceReleased
28 May 2026
Made byVendor
StepFun
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as stepfun/step-3.7-flash.

Route it as stepfun/step-3.7-flash at $0.20 in / $1.15 out per 1M tokens, 262K context. Listed since 28 May 2026.

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p

Run it yourself

Self-hosting

On your own servers,Licence, sizeunder your own control.and hardware.

Compare all open models →All open-weights models →

LicenceLicence
Apache-2.0
OK for business useCommercial use
Yes
SizeParams (total / active)
198B / 11B
Can be trained furtherBase model
No

Hardware you'd need

Hardware tier & memory

Needs one full AI server.

≤ 1.1 TB at 8-bit (one 8×H200 node). Weights ≈ 208 GB at 8-bit, 109 GB at 4-bit (+10–30% for KV cache). MoE: 11B active per token.

Languages: Languages: HF metadata: en only.

Training it further: Fine-tuning: Card: supported in NVIDIA NeMo AutoModel, Megatron Core and Megatron Bridge.

Quantisations: bf16, fp8 (stepfun-ai/Step-3.7-Flash-FP8), nvfp4 (stepfun-ai/Step-3.7-Flash-NVFP4), gguf (stepfun-ai/Step-3.7-Flash-GGUF). Engines: vLLM, SGLang, Transformers, llama.cpp.

Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

No results yetNo results yet

Test results for this model appear after our next daily check.Benchmark results for this model appear after the next daily run.