MiMo-V2.6-Flash
MiMo-V2.6-Flash is Xiaomi's efficiency-tier omnimodal model: 309B MoE with 15B active, 1M context, MIT-licensed weights, and agent, 3D and computer-use
Model Overview
Capabilities, design details, and architectural traits
MiMo-V2.6-Flash - Xiaomi's efficiency-balanced checkpoint from a livestreamed RL run
MiMo-V2.6-Flash is the efficiency tier of Xiaomi's MiMo-V2.6 series, a pair of natively omnimodal models built on the RSI (recursive self-improvement) path: scaling reinforcement learning compute on verifiable, complex tasks so the model keeps expanding its capability frontier through exploration and feedback. Despite being the smaller sibling of MiMo-V2.6-Pro, it comprehensively outperforms the previous generation's flagship, MiMo-V2.5-Pro.
| Differentiator | What it means |
|---|---|
| Efficiency-balanced MoE | 309B total parameters with only 15B activated per token; outperforms the previous MiMo-V2.5-Pro flagship at far lower inference cost |
| Live RL training | 30 RL steps over roughly 750k trajectories in under 6 days, streamed publicly as it happened; held-out DeepSWE v1.1 rose from 48.8 to 65.7 during the run |
| Native omnimodality | Text, image, video and audio enter one model through a MiMo ViT encoder and audio tokenizers, with a 1M-token context |
| Vibe World skills | 3D spatial reasoning, Blender 3D modeling, computer use agent operation and embodied control of a Franka Panda robotic arm |
| Open RL lineage | Published as MiMo-V2.6-Flash-RL under MIT license, with the technical report, training environment and RL code open-sourced |
Scaling RL for self-improvement
The training run expanded RL compute along three documented axes: 1,568 samples per update in a fully asynchronous architecture, a multi-task system spanning Code, General, Visual and Cyber, and greater grader compute that uses within-group comparison to reward shorter paths and fewer tokens. Xiaomi froze the MoE Router to suppress expert load drift and built reward-hacking defenses covering reward design, adversarial evaluation, anomaly detection and validator cross-verification.
From Vibe Coding to Vibe World
MiMo-V2.6-Flash extends natural-language programming into interactive world construction. Given images, video or text, it decomposes game development into multi-agent tasks - 3D scene construction, interactive logic programming and visual verification - then corrects itself from rendering results until it produces a runnable interactive world aligned with user intent.
Benchmark Performance
Independent evaluations · Artificial Analysis
Accuracy & Capability Details
Compare Models Side-by-Side
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | Xiaomi | Anthropic | Anthropic |
| Release Date | September 21, 2026 | September 22, 2026 | September 28, 2026 |
| Knowledge Cutoff | - | - | Jun 2026 |
| Context & Limits | |||
| Context Window | 1M | 1M | 1M |
| Pricing (per 1M tokens) | |||
| Input Pricing | $0.14 Best Input Pricing | $4 | $2 |
| Output Pricing | $0.28 Best Output Pricing | $20 | $10 |
| Modalities | |||
| Inputs | textimageaudiovideo | textimagefile | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 37.9 | 57.6 Best Intelligence Index | 56.0 |
| Coding Index | - | - | - |
| Agentic Index | - | - | - |
Humanity's Last Exam
Extremely difficult logical reasoning and knowledge.
Long Context Reasoning
Logical reasoning over long context windows.
SciCode Benchmark
Scientific coding and mathematical modeling.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.
Explore more from Xiaomi
Other models by Xiaomi
Top AI Models
Leading alternatives by intelligence score