Step 3.7 Flash
Step 3.7 Flash by StepFun is a 198B sparse MoE vision-language model with 11B active params, selectable reasoning levels, and native GUI understanding.
Model Overview
Capabilities, design details, and architectural traits
Step 3.7 Flash – Sparse MoE Vision-Language Model for Agentic Workflows
Step 3.7 Flash is a 198B-parameter sparse Mixture-of-Experts vision-language model from StepFun, built on the Step 3.5 Flash language backbone with a dedicated vision encoder added for native multimodal understanding. Its defining design target is high-frequency production agentic workloads that combine perception, search, and reasoning in a single model without requiring a separate vision module inside agent frameworks.
Multimodal Architecture
Step 3.5 Flash was text-only. Step 3.7 Flash introduces native image understanding by pairing the same language backbone with a 1.8B ViT encoder a structural addition rather than a fine-tune. StepFun documents an observed emergent behavior during testing: the model combined visual tools with non-visual tools without being explicitly trained to do so, such as rendering generated frontend code in a GUI and inspecting the result before iterating.
Visual Search Design
For recognition tasks where parametric knowledge is insufficient such as long-tail entities or recently emerged concepts the model invokes a visual search tool to retrieve and verify. Search is integrated into the reasoning loop rather than treated as a separate add-on.
Deployment
Supports vLLM, SGLang, Hugging Face Transformers, and llama.cpp for inference. Available as an NVIDIA NIM microservice. Local deployment requires at least 128GB unified memory (NVIDIA DGX Station, AMD Ryzen AI Max+ 395, Mac Studio / MacBook Pro).
Benchmark Performance
Independent evaluations · Artificial Analysis
Accuracy & Capability Details
Compare Models Side-by-Side
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | StepFun | Anthropic | Anthropic |
| Release Date | May 29, 2026 | September 22, 2026 | September 28, 2026 |
| Knowledge Cutoff | - | - | Jun 2026 |
| Context & Limits | |||
| Context Window | 256K | 1M Best Context Window | 1M Best Context Window |
| Pricing (per 1M tokens) | |||
| Input Pricing | $0.20 Best Input Pricing | $4 | $2 |
| Output Pricing | $1.15 Best Output Pricing | $20 | $10 |
| Modalities | |||
| Inputs | textimagevideo | textimagefile | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 19.5 | 57.6 Best Intelligence Index | 56.0 |
| Coding Index | 39.6 | - | - |
| Agentic Index | 21.7 | - | - |
Humanity's Last Exam
Extremely difficult logical reasoning and knowledge.
Long Context Reasoning
Logical reasoning over long context windows.
SciCode Benchmark
Scientific coding and mathematical modeling.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.
Explore more from StepFun
Other models by StepFun
Top AI Models
Leading alternatives by intelligence score