StepFun
Released May 29, 2026

Step 3.7 Flash

Step 3.7 Flash by StepFun is a 198B sparse MoE vision-language model with 11B active params, selectable reasoning levels, and native GUI understanding.

Inputs
Text
Image
Video
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Step 3.7 Flash – Sparse MoE Vision-Language Model for Agentic Workflows

Step 3.7 Flash is a 198B-parameter sparse Mixture-of-Experts vision-language model from StepFun, built on the Step 3.5 Flash language backbone with a dedicated vision encoder added for native multimodal understanding. Its defining design target is high-frequency production agentic workloads that combine perception, search, and reasoning in a single model without requiring a separate vision module inside agent frameworks.

Multimodal Architecture

Step 3.5 Flash was text-only. Step 3.7 Flash introduces native image understanding by pairing the same language backbone with a 1.8B ViT encoder a structural addition rather than a fine-tune. StepFun documents an observed emergent behavior during testing: the model combined visual tools with non-visual tools without being explicitly trained to do so, such as rendering generated frontend code in a GUI and inspecting the result before iterating.

Visual Search Design

For recognition tasks where parametric knowledge is insufficient such as long-tail entities or recently emerged concepts the model invokes a visual search tool to retrieve and verify. Search is integrated into the reasoning loop rather than treated as a separate add-on.

Deployment

Supports vLLM, SGLang, Hugging Face Transformers, and llama.cpp for inference. Available as an NVIDIA NIM microservice. Local deployment requires at least 128GB unified memory (NVIDIA DGX Station, AMD Ryzen AI Max+ 395, Mac Studio / MacBook Pro).

Benchmark Performance

Independent evaluations · Artificial Analysis

19.5%
Intelligence
39.6%
Coding Index
21.7%
Agentic Index

Accuracy & Capability Details

GPQA - Graduate Science80.9%
Humanity's Last Exam21.4%
SciCode - Scientific Coding43.9%
Instruction Following67.3%
Long Context Reasoning73.7%
τ²-Bench - Agentic Tasks98.5%
TerminalBench - System Control35.6%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderStepFunAnthropicAnthropic
Release DateMay 29, 2026September 22, 2026September 28, 2026
Knowledge Cutoff--Jun 2026
Context & Limits
Context Window256K
1M
Best Context Window
1M
Best Context Window
Pricing (per 1M tokens)
Input Pricing
$0.20
Best Input Pricing
$4$2
Output Pricing
$1.15
Best Output Pricing
$20$10
Modalities
Inputs
textimagevideo
textimagefile
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index19.5
57.6
Best Intelligence Index
56.0
Coding Index39.6--
Agentic Index21.7--
Step 3.7 Flash
Claude Opus 5.5
Claude Sonnet 5.5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

21%
Step 3.7 Flash
Humanity's Last Exam
Score: 21%
Step 3.7 Flash
61%
Claude Opus 5.5
Humanity's Last Exam
Score: 61%
Claude Opus 5.5
55%
Claude Sonnet 5.5
Humanity's Last Exam
Score: 55%
Claude Sonnet 5.5

Long Context Reasoning

Logical reasoning over long context windows.

74%
Step 3.7 Flash
Long Context Reasoning
Score: 74%
Step 3.7 Flash
85%
Claude Opus 5.5
Long Context Reasoning
Score: 85%
Claude Opus 5.5
83%
Claude Sonnet 5.5
Long Context Reasoning
Score: 83%
Claude Sonnet 5.5

SciCode Benchmark

Scientific coding and mathematical modeling.

44%
Step 3.7 Flash
SciCode Benchmark
Score: 44%
Step 3.7 Flash
67%
Claude Opus 5.5
SciCode Benchmark
Score: 67%
Claude Opus 5.5
61%
Claude Sonnet 5.5
SciCode Benchmark
Score: 61%
Claude Sonnet 5.5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Explore more from StepFun

Other models by StepFun

Top AI Models

Leading alternatives by intelligence score

View all