DeepSeek V4 Flash Vision
DeepSeek V4 Flash Vision is DeepSeek's experimental multimodal model, matching V4-Flash on text while adding image understanding for multimodal agent
Model Overview
Capabilities, design details, and architectural traits
DeepSeek V4 Flash Vision - DeepSeek's Experimental Multimodal Agent Model
DeepSeek V4 Flash Vision is the experimental vision variant of DeepSeek-V4-Flash, accessed as deepseek-v4-flash-vision-exp. Its defining idea is to add image understanding to the Flash model without giving up anything on text. In pure-text work it matches DeepSeek-V4-Flash on agents, reasoning and world knowledge.
| Trait | Detail |
|---|---|
| Experimental status | Released as deepseek-v4-flash-vision-exp, an explicitly experimental model on the DeepSeek API platform |
| Text parity | Matches DeepSeek-V4-Flash on text capabilities, including agents, reasoning and world knowledge |
| Image input methods | Base64 data URLs, external http(s) URLs, or files uploaded once via the Files API and referenced by file_id |
| Image formats | JPEG, PNG, GIF and WebP, with the format detected from actual file content rather than file name or declared MIME type |
| Detail control | low downscales images to 512x512 for faster processing; original keeps the full image |
| API surface | Supports Chat Completions, Messages and Responses APIs, and works across agent frameworks |
Vision Built for Agent Workflows
The model combines visual understanding with a wide range of tools to unlock practical multimodal workflows, such as describing pictures, reading text from screenshots and analyzing charts. It ships with out-of-the-box support in DeepSeek Harness 0.1.1, and the free Files API lets you upload an image once and reuse it across requests to save bandwidth.
Benchmark Performance
Independent evaluations · Artificial Analysis
Accuracy & Capability Details
Compare Models Side-by-Side
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | DeepSeek | Anthropic | Anthropic |
| Release Date | August 21, 2026 | September 22, 2026 | September 28, 2026 |
| Knowledge Cutoff | - | - | Jun 2026 |
| Context & Limits | |||
| Context Window | 1M | 1M | 1M |
| Pricing (per 1M tokens) | |||
| Input Pricing | $0.44 Best Input Pricing | $4 | $2 |
| Output Pricing | $1.32 Best Output Pricing | $20 | $10 |
| Modalities | |||
| Inputs | textimage | textimagefile | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 34.8 | 57.6 Best Intelligence Index | 56.0 |
| Coding Index | 65.0 | - | - |
| Agentic Index | 47.5 | - | - |
Humanity's Last Exam
Extremely difficult logical reasoning and knowledge.
Long Context Reasoning
Logical reasoning over long context windows.
SciCode Benchmark
Scientific coding and mathematical modeling.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.
Explore more from DeepSeek
Other models by DeepSeek
DeepSeek V4.1 Flash
DeepSeek V4 Pro
DeepSeek V4 Flash
DeepSeek V3.2
Top AI Models
Leading alternatives by intelligence score