Ling 3.1 Flash
Ling 3.1 Flash is InclusionAI's efficiency-focused 560B MoE reasoning model with 25B active parameters, 256K context and an explicit thinking mode for agents.
Model Overview
Capabilities, design details, and architectural traits
Ling 3.1 Flash - InclusionAI's efficiency-focused MoE successor
Ling 3.1 Flash is the next-generation efficiency-focused Mixture-of-Experts model from Ant Group's inclusionAI lab, positioned as the successor to Ling-3.0-flash. Its defining move is a large scale jump while keeping the active-parameter footprint small, paired with an explicit thinking mode aimed at agent work.
| Trait | Detail |
|---|---|
| Architecture | Mixture-of-Experts |
| Reasoning | Reasoning model with an explicit thinking mode |
| Target use cases | Agent tasks, search, office software, and specialist applications |
| Context and output | 262,144-token context at launch (1M cited as target), up to 32,768 output tokens; text input and output |
| Launch access | Trial via Vercel's AI Gateway, hosted by Novita, with open weights promised after the trial |
Positioning in the Ling line
The model keeps the Flash positioning of the Ling family: a large-total, small-active MoE designed for efficient serving. The scale jump marks the biggest single-generation change in the line so far.
Release status
At launch the model was available only through the hosted trial. Weights and license were announced as forthcoming but not yet posted, so the model was not yet self-hostable.
Benchmark Performance
Independent evaluations · Artificial Analysis
Accuracy & Capability Details
Compare Models Side-by-Side
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | InclusionAI | Anthropic | Anthropic |
| Release Date | October 1, 2026 | September 22, 2026 | September 28, 2026 |
| Knowledge Cutoff | - | - | Jun 2026 |
| Context & Limits | |||
| Context Window | 262K | 1M Best Context Window | 1M Best Context Window |
| Pricing (per 1M tokens) | |||
| Input Pricing | $0.30 Best Input Pricing | $4 | $2 |
| Output Pricing | $0.90 Best Output Pricing | $20 | $10 |
| Modalities | |||
| Inputs | text | textimagefile | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 41.1 | 57.6 Best Intelligence Index | 56.0 |
| Coding Index | - | - | - |
| Agentic Index | - | - | - |
Humanity's Last Exam
Extremely difficult logical reasoning and knowledge.
Long Context Reasoning
Logical reasoning over long context windows.
SciCode Benchmark
Scientific coding and mathematical modeling.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.
Explore more from InclusionAI
Other models by InclusionAI
Ling-3.0-flash-VL
Ling-3.0-flash-Fin
Ling 3.0 Flash
Ling 3.0 Tiny
Top AI Models
Leading alternatives by intelligence score