Toolbit.aiToolbit.ai

Find, compare, and explore the best AI tools to match your specific tasks and use cases.

Explore

  • AI Search
  • Compare ToolsNew
  • Browse Categories
  • Trending Tools
  • Most Popular
  • New Additions

Resources

  • Updates HubNew
  • AI News
  • ModelsNew
  • Blog Articles
  • NewsletterNew

Company

  • Launch a Tool
  • Advertise with Us
  • Guest Post
  • Contact Us
© 2026 Toolbit.ai. All rights reserved.
Privacy PolicyTerms & ConditionsDisclaimer
There's An AI For That favicon
There's An AI For That•The front page of AI for everyone
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Toolbit.ai
Toolbit.ai
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Sign in
  1. Home
  2. Updates
  3. Models
  4. DeepSeek V3.2 Exp
DeepSeek
Released September 29, 2025Cutoff July 2025

DeepSeek V3.2 Exp

DeepSeek V3.2 Exp by DeepSeek is an experimental 671B MoE model introducing DeepSeek Sparse Attention for validated long-context efficiency gains. MIT licensed, open weights.

Visit DeepSeekAnnouncement
Inputs
Text
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

DeepSeek V3.2 Exp – Experimental Sparse Attention Architecture Validation Model

DeepSeek V3.2 Exp is DeepSeek's experimental release positioned explicitly as an intermediate step toward the next-generation architecture. It builds upon DeepSeek V3.1-Terminus by introducing DeepSeek Sparse Attention (DSA) — a fine-grained sparse attention mechanism designed to explore and validate long-context training and inference efficiency optimizations. The production V3.2 and V3.2-Speciale models share the same model structure as V3.2 Exp.

Architecture

V3.2 Exp shares the same underlying MoE structure as DeepSeek V3.2, with DeepSeek Sparse Attention (DSA) as the defining addition:

  • DSA achieves fine-grained sparse attention — described officially as the first time fine-grained sparse attention of this kind was achieved — delivering substantial improvements in long-context training and inference efficiency while maintaining virtually identical output quality
  • Uses a non-standard sparse attention variant requiring custom code for inference, with custom kernels released across three official repositories: TileLang (research-purpose kernels), DeepGEMM (high-performance CUDA indexer logit kernels including paged versions), and FlashMLA (sparse attention kernels)
  • A documented implementation discrepancy in RoPE within the indexer module was identified and patched post-release: the indexer module requires a non-interleaved layout, while the MLA module expects an interleaved layout.

Benchmark Performance

Independent evaluations · Artificial Analysis

25.9%
Intelligence

Accuracy & Capability Details

GPQA - Graduate Science79.7%
Humanity's Last Exam14.9%
SciCode - Scientific Coding37.7%
Instruction Following54.1%
Long Context Reasoning72.3%
τ²-Bench - Agentic Tasks33.9%
TerminalBench - System Control31.1%
Specs
Context window
164Ktokens
Input pricing
$0.28per 1M tokens
Output pricing
$0.42per 1M tokens
Cached input
$0.03per 1M tokens

Prices in USD.

Key Capabilities & Ratings

Factuality100%
Legal & E-Discovery80%
Scientific Biology80%
Financial Analysis80%
General Knowledge80%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderDeepSeekAnthropicAnthropic
Release DateSeptember 29, 2025July 24, 2026June 9, 2026
Knowledge CutoffJul 2025May 2026-
Context & Limits
Context Window164K
1M
Best Context Window
1M
Best Context Window
Pricing (per 1M tokens)
Input Pricing
$0.28
Best Input Pricing
$5$10
Output Pricing
$0.42
Best Output Pricing
$25$50
Modalities
Inputs
text
textimage
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index21.7
63.1
Best Intelligence Index
62.1
Coding Index-
78.0
Best Coding Index
76.5
Agentic Index-
59.2
Best Agentic Index
56.6
DeepSeek V3.2 Exp
Claude Opus 5
Claude Fable 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

74%
DeepSeek V3.2 Exp
GPQA Benchmark
Score: 74%
DeepSeek V3.2 Exp
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5
93%
Claude Fable 5
GPQA Benchmark
Score: 93%
Claude Fable 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

9%
DeepSeek V3.2 Exp
Humanity's Last Exam
Score: 9%
DeepSeek V3.2 Exp
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5
56%
Claude Fable 5
Humanity's Last Exam
Score: 56%
Claude Fable 5

Long Context Reasoning

Logical reasoning over long context windows.

45%
DeepSeek V3.2 Exp
Long Context Reasoning
Score: 45%
DeepSeek V3.2 Exp
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5
77%
Claude Fable 5
Long Context Reasoning
Score: 77%
Claude Fable 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Back to all models