Kimi K3 by Kimi: first open 3T-class model at 2.8T parameters. Built on Kimi Delta Attention, Attention Residuals, 896-expert MoE, native vision, 1M context.
Capabilities, design details, and architectural traits
Kimi K3 is the first open-source model to reach 2.8 trillion parameters, designed for long-horizon coding, knowledge work, and reasoning. It combines native visual understanding with a 1-million-token context window and an always-on thinking mode.
| Trait | Detail |
|---|---|
| Kimi Delta Attention (KDA) | Hybrid linear attention mechanism that replaces standard quadratic attention in a subset of layers, reducing computational cost across the 1M-token context while preserving expressiveness in critical layers |
| Attention Residuals (AttnRes) | Replaces standard residual connections by allowing each layer to selectively retrieve representations from arbitrary earlier layers, particularly impactful in MoE architectures where different experts activate at different depths |
| Stable LatentMoE | Mixture of Experts framework with 896 experts and 16 active per token, using latent-space routing and Quantile Balancing for load management |
| MXFP4 quantization-aware training | Weights trained in 4-bit floating point (MXFP4) with 8-bit activations (MXFP8) from the supervised fine-tuning stage onward, not post-training quantization, reducing weight storage to ~1.4 TB |
| Always-on thinking mode | Thinking is always enabled with configurable reasoning effort levels: low, high, and max (default max) |
| Native vision | Visual understanding is built into the architecture rather than added via an adapter, enabling screenshot-driven workflows in game development, frontend engineering, and CAD |
Kimi K3 sustains long engineering sessions with minimal human oversight, navigating large codebases and orchestrating terminal tools. It combines software engineering with visual reasoning, using screenshots and visual feedback to improve workflows.
The combination of KDA, AttnRes, increased MoE sparsity, and refined training recipes yields approximately 2.5x the overall scaling efficiency of the predecessor Kimi K2, converting compute into capability more effectively.
Independent evaluations · Artificial Analysis
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | Kimi | Anthropic | Anthropic |
| Release Date | July 16, 2026 | July 24, 2026 | June 9, 2026 |
| Knowledge Cutoff | - | May 2026 | - |
| Context & Limits | |||
| Context Window | 1M | 1M | 1M |
| Pricing (per 1M tokens) | |||
| Input Pricing | $3 Best Input Pricing | $5 | $10 |
| Output Pricing | $15 Best Output Pricing | $25 | $50 |
| Modalities | |||
| Inputs | text | textimage | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 59.7 | 63.1 Best Intelligence Index | 62.1 |
| Coding Index | 76.2 | 78.0 Best Coding Index | 76.5 |
| Agentic Index | 54.3 | 59.2 Best Agentic Index | 56.6 |
Graduate-level reasoning and expert Q&A evaluation.
Extremely difficult logical reasoning and knowledge.
Logical reasoning over long context windows.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.