Toolbit.aiToolbit.ai

Find, compare, and explore the best AI tools to match your specific tasks and use cases.

Explore

  • AI Search
  • Compare ToolsNew
  • Browse Categories
  • Trending Tools
  • Most Popular
  • New Additions

Resources

  • Updates HubNew
  • AI News
  • ModelsNew
  • Blog Articles
  • NewsletterNew

Company

  • Launch a Tool
  • Advertise with Us
  • Guest Post
  • Contact Us
© 2026 Toolbit.ai. All rights reserved.
Privacy PolicyTerms & ConditionsDisclaimer
There's An AI For That favicon
There's An AI For That•The front page of AI for everyone
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Toolbit.ai
Toolbit.ai
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Sign in
  1. Home
  2. Updates
  3. Models
  4. DeepSeek V3 0324
DeepSeek
Released March 25, 2025Cutoff July 2024

DeepSeek V3 0324

DeepSeek V3 0324 by DeepSeek is a March 2025 update to DeepSeek V3, sharing identical architecture with improved post-training for reasoning, coding, and function calling.

Visit DeepSeekAnnouncement
Inputs
Text
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

DeepSeek V3 0324 – Post-Training-Updated Checkpoint of DeepSeek V3

DeepSeek V3 0324 is a March 25, 2025 update to DeepSeek V3, released as an open-weight checkpoint with the same underlying model structure as the original V3. The defining change is an improved post-training pipeline that draws on RL techniques from DeepSeek R1, producing measurable gains in reasoning, coding, front-end web development, and function calling without any architectural modification.

Architecture

Officially documented as identical in model structure to DeepSeek V3: a 671B total parameter, 37B active parameter Mixture-of-Experts model using Multi-head Latent Attention (MLA), DeepSeekMoE, auxiliary-loss-free load balancing, and Multi-Token Prediction (MTP). No structural changes were introduced in this update.

Post-Training Changes

The officially documented improvements over DeepSeek V3 are:

  • Reasoning: Post-training pipeline updated drawing on R1 RL techniques; AIME score improved from 39.6 → 59.4 (+19.8), GPQA from 59.1 → 68.4 (+9.3), MMLU-Pro from 75.9 → 81.2 (+5.3)
  • Coding: LiveCodeBench improved from 39.2 → 49.2 (+10.0); improved accuracy in code generation; more aesthetically refined front-end web and game output
  • Chinese writing: Style and content quality aligned with R1 writing style; improved medium-to-long-form writing
  • Function Calling: Increased accuracy; fixes to issues present in the previous V3 version
  • Translation and multi-turn rewriting: Optimized quality.

Benchmark Performance

Independent evaluations · Artificial Analysis

15.2%
Intelligence
21.2%
Coding Index
1.6%
Agentic Index

Accuracy & Capability Details

GPQA - Graduate Science65.5%
Humanity's Last Exam4.7%
SciCode - Scientific Coding35.8%
Instruction Following41.0%
Long Context Reasoning41.3%
τ²-Bench - Agentic Tasks47.1%
TerminalBench - System Control15.2%
Specs
Context window
164Ktokens
Input pricing
$0.27per 1M tokens
Output pricing
$1.12per 1M tokens
Cached input
$0.27per 1M tokens

Prices in USD.

Key Capabilities & Ratings

Advanced Math80%
Legal & E-Discovery80%
Financial Analysis80%
Translation & Lang80%
Medical Reasoning80%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderDeepSeekAnthropicAnthropic
Release DateMarch 25, 2025July 24, 2026June 9, 2026
Knowledge CutoffJul 2024May 2026-
Context & Limits
Context Window164K
1M
Best Context Window
1M
Best Context Window
Pricing (per 1M tokens)
Input Pricing
$0.27
Best Input Pricing
$5$10
Output Pricing
$1.12
Best Output Pricing
$25$50
Modalities
Inputs
text
textimage
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index15.2
63.1
Best Intelligence Index
62.1
Coding Index21.2
78.0
Best Coding Index
76.5
Agentic Index1.6
59.2
Best Agentic Index
56.6
DeepSeek V3 0324
Claude Opus 5
Claude Fable 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

66%
DeepSeek V3 0324
GPQA Benchmark
Score: 66%
DeepSeek V3 0324
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5
93%
Claude Fable 5
GPQA Benchmark
Score: 93%
Claude Fable 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

5%
DeepSeek V3 0324
Humanity's Last Exam
Score: 5%
DeepSeek V3 0324
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5
56%
Claude Fable 5
Humanity's Last Exam
Score: 56%
Claude Fable 5

Long Context Reasoning

Logical reasoning over long context windows.

41%
DeepSeek V3 0324
Long Context Reasoning
Score: 41%
DeepSeek V3 0324
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5
77%
Claude Fable 5
Long Context Reasoning
Score: 77%
Claude Fable 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Back to all models