DeepSeek V3 by DeepSeek is a 671B/37B-active MoE model with Multi-head Latent Attention, auxiliary-loss-free load balancing, and Multi-Token Prediction. MIT licensed.
Capabilities, design details, and architectural traits
DeepSeek V3 is DeepSeek's flagship general-purpose model, built around the co-design of three documented architectural and training innovations: Multi-head Latent Attention (MLA) for inference efficiency, DeepSeekMoE for cost-effective training, and a pioneering auxiliary-loss-free load balancing strategy. The full training required 2.788M H800 GPU hours with no irrecoverable loss spikes or rollbacks throughout.
DeepSeek V3 is a 671B total parameter, 37B active parameter Mixture-of-Experts model. Two core architectural components carry over from DeepSeek V2 and are the structural foundation:
Two additional strategies are introduced in V3:
Independent evaluations · Artificial Analysis
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | DeepSeek | Anthropic | Anthropic |
| Release Date | March 25, 2025 | July 24, 2026 | June 9, 2026 |
| Knowledge Cutoff | Jul 2024 | May 2026 | - |
| Context & Limits | |||
| Context Window | 131K | 1M Best Context Window | 1M Best Context Window |
| Pricing (per 1M tokens) | |||
| Input Pricing | $0.36 Best Input Pricing | $5 | $10 |
| Output Pricing | $0.89 Best Output Pricing | $25 | $50 |
| Modalities | |||
| Inputs | text | textimage | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 14.2 | 63.1 Best Intelligence Index | 62.1 |
| Coding Index | 23.0 | 78.0 Best Coding Index | 76.5 |
| Agentic Index | 1.6 | 59.2 Best Agentic Index | 56.6 |
Graduate-level reasoning and expert Q&A evaluation.
Extremely difficult logical reasoning and knowledge.
Logical reasoning over long context windows.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.