DeepSeek V4.1 Flash
DeepSeek V4.1 Flash is a 552B multimodal model with a 196B-parameter Engram memory system, native vision, 1M context, and one-eighth V4-Flash's KV cache.
Model Overview
Capabilities, design details, and architectural traits
DeepSeek V4.1 Flash - the Engram memory model built for long-running agents
DeepSeek V4.1 Flash is a multimodal model designed around a specific problem: the longer an AI agent works, the more conversation history it must carry, and the more expensive it becomes to keep alive. Rather than shrinking the model, DeepSeek redesigned how much of it has to work, how much previous context it stores, and how often it can reuse computation it already performed.
The result is a large model paired with a dedicated memory architecture and aggressive cache reduction aimed at sustained agent workloads.
What sets V4.1 Flash apart
| Trait | Detail |
|---|---|
| Engram memory system | A dedicated memory subsystem, separate from the backbone |
| Persistent KV-cache reduction | Roughly one-eighth the persistent KV-cache storage of V4-Flash at the same sequence length |
| Global KV-cache reduction | About one-quarter the global KV cache of V4-Flash |
| Native vision | Built-in image understanding rather than a bolted-on vision adapter |
| 1M token context | A context window of up to one million tokens |
| Agent-first design | Optimized for multi-hour tool-using workflows where history grows with every step |
Why memory, not size, is the story
The technical report behind V4.1 Flash frames persistent memory cost as the core constraint for agents that research, open files, run code, and check their own work over long sessions. The Engram system and the compressed KV cache are the mechanisms that let the model stay practical for these workloads.
The model is positioned as the efficiency tier of the V4.1 generation: a compact, native-vision model delivering faster inference and higher throughput, with open weights and a technical report published on Hugging Face.
Benchmark Performance
Independent evaluations · Artificial Analysis
Accuracy & Capability Details
Compare Models Side-by-Side
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | DeepSeek | Anthropic | Anthropic |
| Release Date | September 10, 2026 | September 22, 2026 | September 28, 2026 |
| Knowledge Cutoff | - | - | Jun 2026 |
| Context & Limits | |||
| Context Window | 1M | 1M | 1M |
| Pricing (per 1M tokens) | |||
| Input Pricing | $0.30 Best Input Pricing | $4 | $2 |
| Output Pricing | $1.20 Best Output Pricing | $20 | $10 |
| Modalities | |||
| Inputs | textimage | textimagefile | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 39.5 | 57.6 Best Intelligence Index | 56.0 |
| Coding Index | - | - | - |
| Agentic Index | - | - | - |
Humanity's Last Exam
Extremely difficult logical reasoning and knowledge.
Long Context Reasoning
Logical reasoning over long context windows.
SciCode Benchmark
Scientific coding and mathematical modeling.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.
Explore more from DeepSeek
Other models by DeepSeek
DeepSeek V4 Flash Vision
DeepSeek V4 Pro
DeepSeek V4 Flash
DeepSeek V3.2
Top AI Models
Leading alternatives by intelligence score