DeepSeek V4 Pro
DeepSeek V4 Pro by DeepSeek is a 1.6T/49B-active MoE model with hybrid CSA+HCA attention, three reasoning modes, and 1M-token context. Open weights.
Model Overview
Capabilities, design details, and architectural traits
DeepSeek V4 Pro - 1.6T Sparse MoE with Hybrid Long-Context Attention
DeepSeek V4 Pro is DeepSeek's flagship model in the V4 series: a 1.6 trillion total parameter, 49B active parameter Mixture-of-Experts language model built around a newly designed hybrid attention architecture engineered specifically to make 1 million-token context windows computationally viable.
Architecture
The V4 series introduces three documented architectural innovations over prior DeepSeek generations:
- Hybrid Attention (CSA + HCA): Combines Compressed Sparse Attention and Heavily Compressed Attention across layers to reduce long-context inference cost without degrading comprehension
- Manifold-Constrained Hyper-Connections (mHC): Strengthens residual connections to stabilize signal propagation across the model's deep layer stack while preserving expressivity
- Mixed Precision Training: MoE expert parameters use FP4 precision; most other parameters use FP8
Reasoning Workflow
Three configurable reasoning effort modes per request:
- Non-think - fast responses, no chain-of-thought
- Think High - logical analysis with reasoning
- Think Max (V4-Pro-Max) - maximum reasoning effort; recommended minimum 384K context
Reasoning is surfaced via <think> tags in output.
Benchmark Performance
Independent evaluations · Artificial Analysis
Accuracy & Capability Details
Compare Models Side-by-Side
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | DeepSeek | Anthropic | Anthropic |
| Release Date | April 24, 2026 | September 22, 2026 | September 28, 2026 |
| Knowledge Cutoff | - | - | Jun 2026 |
| Context & Limits | |||
| Context Window | 1.0M Best Context Window | 1M | 1M |
| Pricing (per 1M tokens) | |||
| Input Pricing | $0.43 Best Input Pricing | $4 | $2 |
| Output Pricing | $0.87 Best Output Pricing | $20 | $10 |
| Modalities | |||
| Inputs | text | textimagefile | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 30.1 | 57.6 Best Intelligence Index | 56.0 |
| Coding Index | 58.7 | - | - |
| Agentic Index | 35.3 | - | - |
Humanity's Last Exam
Extremely difficult logical reasoning and knowledge.
Long Context Reasoning
Logical reasoning over long context windows.
SciCode Benchmark
Scientific coding and mathematical modeling.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.
Explore more from DeepSeek
Other models by DeepSeek
DeepSeek V4.1 Flash
DeepSeek V4 Flash Vision
DeepSeek V4 Flash
DeepSeek V3.2
Top AI Models
Leading alternatives by intelligence score