MiMo-V2.6-Pro
MiMo-V2.6-Pro is Xiaomi's trillion-parameter omni-modal reasoning model, trained by scaling RL on verifiable tasks, with a 1M-token context window.
Model Overview
Capabilities, design details, and architectural traits
MiMo-V2.6-Pro - Xiaomi's trillion-parameter omni-modal flagship
MiMo-V2.6-Pro is Xiaomi's flagship reasoning model. It is natively omni-modal, jointly understanding images, video, audio, and text, and it is fully open-sourced. It is built for complex projects, long-horizon tasks, high-stakes work, cybersecurity, and research.
| Trait | Detail |
|---|---|
| Creator | Xiaomi |
| Identity | Trillion-parameter, omni-modal flagship reasoning model, released open-source |
| Training approach | RSI path: scaling RL compute on verifiable, complex tasks so the model expands its capability frontier through exploration and feedback |
| RL run | 30 RL steps over roughly 750k trajectories in under six days, with the production run streamed live as it happened |
| Context window | 1M tokens input, 128K tokens max output |
| Capabilities | Deep thinking, tool call, streaming, web search, structured output, context caching |
| Collaboration style | Tuned for multimodal, multi-agent, and multi-harness collaboration; breaks tasks into sub-tasks and coordinates multiple agents |
| Target workflows | Coding, office work, design, research, content creation, cybersecurity, and computer operation |
Built in public through RL
What separates MiMo-V2.6-Pro from most flagship releases is how it was trained. Xiaomi describes its approach as the RSI path, scaling reinforcement learning compute on verifiable, complex tasks rather than relying on standard post-training alone. The full production run was streamed live, and Xiaomi reports that RL proved sample-efficient and generalized beyond the training distribution.
From coding to co-scientist
Xiaomi highlights workflows where the model acts as a coordinator rather than a single responder. In its 3D open-world game demo, it decomposes a prompt into sub-tasks and directs multiple agents to build scenes, interaction logic, and visual checks. In materials research, it works as a co-scientist, reviewing literature, forming hypotheses, and simulating binding strength for candidate MOF materials.
Benchmark Performance
Independent evaluations · Artificial Analysis
Accuracy & Capability Details
Compare Models Side-by-Side
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | Xiaomi | Anthropic | Anthropic |
| Release Date | September 21, 2026 | September 22, 2026 | September 1, 2026 |
| Knowledge Cutoff | Dec 2024 | - | - |
| Context & Limits | |||
| Context Window | 1M | 1M | 1M |
| Pricing (per 1M tokens) | |||
| Input Pricing | $0.43 Best Input Pricing | $4 | $10 |
| Output Pricing | $0.87 Best Output Pricing | $20 | $50 |
| Modalities | |||
| Inputs | textimageaudiovideo | textimagefile | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 46.3 | 57.6 Best Intelligence Index | 53.4 |
| Coding Index | - | - | 81.6 |
| Agentic Index | - | - | 57.9 |
Humanity's Last Exam
Extremely difficult logical reasoning and knowledge.
Long Context Reasoning
Logical reasoning over long context windows.
SciCode Benchmark
Scientific coding and mathematical modeling.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.