DeepSeek
Released August 21, 2026

DeepSeek V4 Flash Vision

DeepSeek V4 Flash Vision is DeepSeek's experimental multimodal model, matching V4-Flash on text while adding image understanding for multimodal agent

Inputs
Text
Image
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

DeepSeek V4 Flash Vision - DeepSeek's Experimental Multimodal Agent Model

DeepSeek V4 Flash Vision is the experimental vision variant of DeepSeek-V4-Flash, accessed as deepseek-v4-flash-vision-exp. Its defining idea is to add image understanding to the Flash model without giving up anything on text. In pure-text work it matches DeepSeek-V4-Flash on agents, reasoning and world knowledge.

TraitDetail
Experimental statusReleased as deepseek-v4-flash-vision-exp, an explicitly experimental model on the DeepSeek API platform
Text parityMatches DeepSeek-V4-Flash on text capabilities, including agents, reasoning and world knowledge
Image input methodsBase64 data URLs, external http(s) URLs, or files uploaded once via the Files API and referenced by file_id
Image formatsJPEG, PNG, GIF and WebP, with the format detected from actual file content rather than file name or declared MIME type
Detail controllow downscales images to 512x512 for faster processing; original keeps the full image
API surfaceSupports Chat Completions, Messages and Responses APIs, and works across agent frameworks

Vision Built for Agent Workflows

The model combines visual understanding with a wide range of tools to unlock practical multimodal workflows, such as describing pictures, reading text from screenshots and analyzing charts. It ships with out-of-the-box support in DeepSeek Harness 0.1.1, and the free Files API lets you upload an image once and reuse it across requests to save bandwidth.

Benchmark Performance

Independent evaluations · Artificial Analysis

34.8%
Intelligence
65.0%
Coding Index
47.5%
Agentic Index

Accuracy & Capability Details

GPQA - Graduate Science91.3%
Humanity's Last Exam34.5%
SciCode - Scientific Coding49.7%
Long Context Reasoning81.3%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderDeepSeekAnthropicAnthropic
Release DateAugust 21, 2026September 22, 2026September 28, 2026
Knowledge Cutoff--Jun 2026
Context & Limits
Context Window1M1M1M
Pricing (per 1M tokens)
Input Pricing
$0.44
Best Input Pricing
$4$2
Output Pricing
$1.32
Best Output Pricing
$20$10
Modalities
Inputs
textimage
textimagefile
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index34.8
57.6
Best Intelligence Index
56.0
Coding Index65.0--
Agentic Index47.5--
DeepSeek V4 Flash Vision
Claude Opus 5.5
Claude Sonnet 5.5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

35%
DeepSeek V4 Flash Vision
Humanity's Last Exam
Score: 35%
DeepSeek V4 Flash Vision
61%
Claude Opus 5.5
Humanity's Last Exam
Score: 61%
Claude Opus 5.5
55%
Claude Sonnet 5.5
Humanity's Last Exam
Score: 55%
Claude Sonnet 5.5

Long Context Reasoning

Logical reasoning over long context windows.

81%
DeepSeek V4 Flash Vision
Long Context Reasoning
Score: 81%
DeepSeek V4 Flash Vision
85%
Claude Opus 5.5
Long Context Reasoning
Score: 85%
Claude Opus 5.5
83%
Claude Sonnet 5.5
Long Context Reasoning
Score: 83%
Claude Sonnet 5.5

SciCode Benchmark

Scientific coding and mathematical modeling.

50%
DeepSeek V4 Flash Vision
SciCode Benchmark
Score: 50%
DeepSeek V4 Flash Vision
67%
Claude Opus 5.5
SciCode Benchmark
Score: 67%
Claude Opus 5.5
61%
Claude Sonnet 5.5
SciCode Benchmark
Score: 61%
Claude Sonnet 5.5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Explore more from DeepSeek

Other models by DeepSeek

Top AI Models

Leading alternatives by intelligence score

View all