Toolbit.aiToolbit.ai

Find, compare, and explore the best AI tools to match your specific tasks and use cases.

Explore

  • AI Search
  • Compare ToolsNew
  • Browse Categories
  • Trending Tools
  • Most Popular
  • New Additions

Resources

  • Updates HubNew
  • AI News
  • ModelsNew
  • Blog Articles
  • NewsletterNew

Company

  • Launch a Tool
  • Advertise with Us
  • Guest Post
  • Contact Us
© 2026 Toolbit.ai. All rights reserved.
Privacy PolicyTerms & ConditionsDisclaimer
There's An AI For That favicon
There's An AI For That•The front page of AI for everyone
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Toolbit.ai
Toolbit.ai
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Sign in
  1. Home
  2. Updates
  3. Models
  4. Step 3.5 Flash
StepFun
Released February 2, 2026

Step 3.5 Flash

Step 3.5 Flash by StepFun is an open-source sparse MoE model with 196B total/11B active parameters, MTP-3 multi-token prediction, and 256K context for agentic workflows.

Visit StepFunAnnouncement
Inputs
Text
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Step 3.5 Flash – Sparse MoE Agentic Reasoning Model

Step 3.5 Flash is StepFun's most capable open-source foundation model, built around the concept of intelligence density: a 196B-parameter sparse Mixture-of-Experts architecture that activates only ~11B parameters per token, targeting frontier-level reasoning at the latency profile required for real-time agentic interaction.

Architecture

The model uses a 45-layer sparse-MoE Transformer backbone (3 dense + 42 MoE layers). Each MoE layer contains 288 routed experts plus 1 always-active shared expert, with a top-8 router selecting active experts per token. This configuration preserves large-model knowledge capacity while constraining per-token compute.

Inference Design

Three co-designed mechanisms reduce wall-clock latency specifically for multi-turn agentic decoding:

  • MTP-3 (3-way Multi-Token Prediction): Predicts 4 tokens per forward pass via a specialized head using sliding-window attention and a dense FFN, achieving 100–300 tok/s throughput (peaking at 350 tok/s on single-stream coding tasks)
  • Hybrid Attention (3:1 SWA:Full ratio): Three sliding-window attention layers per full-attention layer, with window size 512, reducing long-context memory overhead
  • Head-wise gated attention: Replaces fixed sink tokens with per-head gating, documented to raise pretraining benchmark averages.

Benchmark Performance

Independent evaluations · Artificial Analysis

26.0%
Intelligence

Accuracy & Capability Details

GPQA - Graduate Science83.1%
Humanity's Last Exam21.1%
SciCode - Scientific Coding40.4%
Instruction Following64.6%
Long Context Reasoning48.0%
τ²-Bench - Agentic Tasks94.4%
TerminalBench - System Control27.3%
Specs
Context window
262Ktokens
Input pricing
$0.10per 1M tokens
Output pricing
$0.30per 1M tokens

Prices in USD.

Key Capabilities & Ratings

Advanced Math90%
General Knowledge90%
Logical Logic80%
Agentic Actions70%
Search Integration70%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderStepFunAnthropicAnthropic
Release DateFebruary 2, 2026July 24, 2026June 9, 2026
Knowledge Cutoff-May 2026-
Context & Limits
Context Window262K
1M
Best Context Window
1M
Best Context Window
Pricing (per 1M tokens)
Input Pricing
$0.10
Best Input Pricing
$5$10
Output Pricing
$0.30
Best Output Pricing
$25$50
Modalities
Inputs
text
textimage
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index26.0
63.1
Best Intelligence Index
62.1
Coding Index-
78.0
Best Coding Index
76.5
Agentic Index-
59.2
Best Agentic Index
56.6
Step 3.5 Flash
Claude Opus 5
Claude Fable 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

83%
Step 3.5 Flash
GPQA Benchmark
Score: 83%
Step 3.5 Flash
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5
93%
Claude Fable 5
GPQA Benchmark
Score: 93%
Claude Fable 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

21%
Step 3.5 Flash
Humanity's Last Exam
Score: 21%
Step 3.5 Flash
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5
56%
Claude Fable 5
Humanity's Last Exam
Score: 56%
Claude Fable 5

Long Context Reasoning

Logical reasoning over long context windows.

48%
Step 3.5 Flash
Long Context Reasoning
Score: 48%
Step 3.5 Flash
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5
77%
Claude Fable 5
Long Context Reasoning
Score: 77%
Claude Fable 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Back to all models