Ling 3.0 Tiny by InclusionAI is a 7.9B MoE model with 1.3B active parameters for responsive agents, multi-turn conversations, and switchable thinking modes.
Capabilities, design details, and architectural traits
Ling 3.0 Tiny is a mixture-of-experts model from InclusionAI that activates only a small fraction of its parameters per token. This unusually low active-compute ratio is the model's defining design choice, enabling it to serve agent and multi-turn conversation workloads with minimal per-token cost.
The model supports switchable Thinking and Instant modes, letting callers choose between deeper reasoning and fast responses within the same deployment. It also includes native function calling and prompt caching for tool-using, long-context workflows.
| Trait | Detail |
|---|---|
| Architecture | MoE with low active-compute ratio |
| Switchable modes | Thinking mode and Instant mode, selectable per request |
| Agent design focus | Built for responsive agents, instruction following, and multi-turn conversations |
| Tool support | Native function calling and prompt caching |
| Context window | 256K tokens with up to 32K max output |
| I/O modality | Text input, text output only |
The switchable Thinking and Instant modes distinguish Ling 3.0 Tiny from single-mode models in its class. Thinking mode engages deeper reasoning for complex prompts, while Instant mode prioritizes low-latency responses for conversational turns where speed matters more than extended deliberation.
Independent evaluations · Artificial Analysis
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | InclusionAI | Anthropic | Anthropic |
| Release Date | August 6, 2026 | July 24, 2026 | June 9, 2026 |
| Knowledge Cutoff | - | May 2026 | - |
| Context & Limits | |||
| Context Window | 262K | 1M Best Context Window | 1M Best Context Window |
| Pricing (per 1M tokens) | |||
| Input Pricing | Free Best Input Pricing | $5 | $10 |
| Output Pricing | Free Best Output Pricing | $25 | $50 |
| Modalities | |||
| Inputs | text | textimage | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 24.5 | 63.1 Best Intelligence Index | 62.1 |
| Coding Index | 26.5 | 78.0 Best Coding Index | 76.5 |
| Agentic Index | 16.0 | 59.2 Best Agentic Index | 56.6 |
Graduate-level reasoning and expert Q&A evaluation.
Extremely difficult logical reasoning and knowledge.
Logical reasoning over long context windows.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.