DeepSeek
Released September 10, 2026

DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is a 552B multimodal model with a 196B-parameter Engram memory system, native vision, 1M context, and one-eighth V4-Flash's KV cache.

Inputs
Text
Image
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

DeepSeek V4.1 Flash - the Engram memory model built for long-running agents

DeepSeek V4.1 Flash is a multimodal model designed around a specific problem: the longer an AI agent works, the more conversation history it must carry, and the more expensive it becomes to keep alive. Rather than shrinking the model, DeepSeek redesigned how much of it has to work, how much previous context it stores, and how often it can reuse computation it already performed.

The result is a large model paired with a dedicated memory architecture and aggressive cache reduction aimed at sustained agent workloads.

What sets V4.1 Flash apart

TraitDetail
Engram memory systemA dedicated memory subsystem, separate from the backbone
Persistent KV-cache reductionRoughly one-eighth the persistent KV-cache storage of V4-Flash at the same sequence length
Global KV-cache reductionAbout one-quarter the global KV cache of V4-Flash
Native visionBuilt-in image understanding rather than a bolted-on vision adapter
1M token contextA context window of up to one million tokens
Agent-first designOptimized for multi-hour tool-using workflows where history grows with every step

Why memory, not size, is the story

The technical report behind V4.1 Flash frames persistent memory cost as the core constraint for agents that research, open files, run code, and check their own work over long sessions. The Engram system and the compressed KV cache are the mechanisms that let the model stay practical for these workloads.

The model is positioned as the efficiency tier of the V4.1 generation: a compact, native-vision model delivering faster inference and higher throughput, with open weights and a technical report published on Hugging Face.

Benchmark Performance

Independent evaluations · Artificial Analysis

39.5%
Intelligence

Accuracy & Capability Details

Humanity's Last Exam39.2%
SciCode - Scientific Coding51.9%
Long Context Reasoning84.0%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderDeepSeekAnthropicAnthropic
Release DateSeptember 10, 2026September 22, 2026September 28, 2026
Knowledge Cutoff--Jun 2026
Context & Limits
Context Window1M1M1M
Pricing (per 1M tokens)
Input Pricing
$0.30
Best Input Pricing
$4$2
Output Pricing
$1.20
Best Output Pricing
$20$10
Modalities
Inputs
textimage
textimagefile
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index39.5
57.6
Best Intelligence Index
56.0
Coding Index---
Agentic Index---
DeepSeek V4.1 Flash
Claude Opus 5.5
Claude Sonnet 5.5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

39%
DeepSeek V4.1 Flash
Humanity's Last Exam
Score: 39%
DeepSeek V4.1 Flash
61%
Claude Opus 5.5
Humanity's Last Exam
Score: 61%
Claude Opus 5.5
55%
Claude Sonnet 5.5
Humanity's Last Exam
Score: 55%
Claude Sonnet 5.5

Long Context Reasoning

Logical reasoning over long context windows.

84%
DeepSeek V4.1 Flash
Long Context Reasoning
Score: 84%
DeepSeek V4.1 Flash
85%
Claude Opus 5.5
Long Context Reasoning
Score: 85%
Claude Opus 5.5
83%
Claude Sonnet 5.5
Long Context Reasoning
Score: 83%
Claude Sonnet 5.5

SciCode Benchmark

Scientific coding and mathematical modeling.

52%
DeepSeek V4.1 Flash
SciCode Benchmark
Score: 52%
DeepSeek V4.1 Flash
67%
Claude Opus 5.5
SciCode Benchmark
Score: 67%
Claude Opus 5.5
61%
Claude Sonnet 5.5
SciCode Benchmark
Score: 61%
Claude Sonnet 5.5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Explore more from DeepSeek

Other models by DeepSeek

Top AI Models

Leading alternatives by intelligence score

View all