Toolbit.aiToolbit.ai

Find, compare, and explore the best AI tools to match your specific tasks and use cases.

Explore

  • AI Search
  • Compare ToolsNew
  • Browse Categories
  • Trending Tools
  • Most Popular
  • New Additions

Resources

  • Updates HubNew
  • AI News
  • ModelsNew
  • Blog Articles
  • NewsletterNew

Company

  • Launch a Tool
  • Advertise with Us
  • Guest Post
  • Contact Us
© 2026 Toolbit.ai. All rights reserved.
Privacy PolicyTerms & ConditionsDisclaimer
There's An AI For That favicon
There's An AI For That•The front page of AI for everyone
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Toolbit.ai
Toolbit.ai
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Sign in
  1. Home
  2. Updates
  3. Models
  4. Ling 3.0 Flash
InclusionAI
Released August 4, 2026

Ling 3.0 Flash

Ling 3.0 Flash by InclusionAI: 124B hybrid-linear MoE, 5.1B active per token. KDA+MLA attention, native hybrid reasoning, 10K+ agentic training environments.

Visit InclusionAIAnnouncement
Inputs
Text
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Ling 3.0 Flash - Native Hybrid-Linear Reasoning MoE

Ling 3.0 Flash is a hybrid-linear Mixture-of-Experts model from InclusionAI.

The model's defining architectural choice is a native hybrid-linear attention stack built from the start of pretraining. It alternates Kimi Delta Attention (KDA) layers with Gated Multi-head Latent Attention (MLA) in a 5:1 ratio (35 KDA + 7 Gated MLA + 2 dense layers). KDA introduces fine-grained diagonal gating in Delta Rule state updates, keeping long-context memory stable and predictable. Periodic full-attention MLA layers preserve exact token-to-token recall that pure linear attention loses.

TraitDetail
Ultra-sparse MoE512 routed experts + 1 shared expert; only 8 experts activated per token (1/64 ratio, down from 1/32 in prior generation)
Native hybrid reasoningCombines Flash-series speed with Ring-series deep thinking; thinking mode enabled by default, dynamically scaling effort by task difficulty
Agentic training10,000+ interactive training environments for end-to-end closed-loop execution across coding, general, and deep research agent tasks
Hierarchical cachingNatively integrates SGLang HiCache + Mooncake architecture with physical dual-pools and cluster-shared L3 cache, reducing TTFT by 60-80% in long-input scenarios
Context scheduleTrained progressively at 8K -> 32K -> 256K context windows

Bridging Flash and Ring

InclusionAI's model families serve distinct roles: Ling models prioritize fast, high-throughput production inference, while Ring models specialize in deep step-by-step reasoning. Ling 3.0 Flash fuses both lineages into a single architecture, dynamically switching between rapid non-thinking responses and multi-step reasoning depending on task complexity. This hybrid reasoning mode is the model's core differentiator against pure throughput or pure reasoning models.

Benchmark Performance

Independent evaluations · Artificial Analysis

37.8%
Intelligence
50.6%
Coding Index
29.3%
Agentic Index

Accuracy & Capability Details

GPQA - Graduate Science85.5%
Humanity's Last Exam23.7%
SciCode - Scientific Coding41.1%
Long Context Reasoning67.0%
Specs
Context window
262Ktokens
Input pricing
$0.07per 1M tokens
Output pricing
$0.22per 1M tokens
Cached input
$0.01per 1M tokens

Prices in USD.

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderInclusionAIAnthropicAnthropic
Release DateAugust 4, 2026July 24, 2026June 9, 2026
Knowledge Cutoff-May 2026-
Context & Limits
Context Window262K
1M
Best Context Window
1M
Best Context Window
Pricing (per 1M tokens)
Input Pricing
$0.07
Best Input Pricing
$5$10
Output Pricing
$0.22
Best Output Pricing
$25$50
Modalities
Inputs
text
textimage
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index37.8
63.1
Best Intelligence Index
62.1
Coding Index50.6
78.0
Best Coding Index
76.5
Agentic Index29.3
59.2
Best Agentic Index
56.6
Ling 3.0 Flash
Claude Opus 5
Claude Fable 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

86%
Ling 3.0 Flash
GPQA Benchmark
Score: 86%
Ling 3.0 Flash
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5
93%
Claude Fable 5
GPQA Benchmark
Score: 93%
Claude Fable 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

24%
Ling 3.0 Flash
Humanity's Last Exam
Score: 24%
Ling 3.0 Flash
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5
56%
Claude Fable 5
Humanity's Last Exam
Score: 56%
Claude Fable 5

Long Context Reasoning

Logical reasoning over long context windows.

67%
Ling 3.0 Flash
Long Context Reasoning
Score: 67%
Ling 3.0 Flash
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5
77%
Claude Fable 5
Long Context Reasoning
Score: 77%
Claude Fable 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Back to all models