Toolbit.aiToolbit.ai

Find, compare, and explore the best AI tools to match your specific tasks and use cases.

Explore

  • AI Search
  • Compare ToolsNew
  • Browse Categories
  • Trending Tools
  • Most Popular
  • New Additions

Resources

  • Updates HubNew
  • AI News
  • ModelsNew
  • Blog Articles
  • NewsletterNew

Company

  • Launch a Tool
  • Advertise with Us
  • Guest Post
  • Contact Us
© 2026 Toolbit.ai. All rights reserved.
Privacy PolicyTerms & ConditionsDisclaimer
There's An AI For That favicon
There's An AI For That•The front page of AI for everyone
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Toolbit.ai
Toolbit.ai
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Sign in
  1. Home
  2. Updates
  3. Models
  4. Claude Sonnet 4
Anthropic
Released May 22, 2025Cutoff January 2025

Claude Sonnet 4

Claude Sonnet 4 by Anthropic is a hybrid reasoning model with extended thinking + tool use, parallel tool calls, enhanced steerability, and free-user access.

Visit AnthropicAnnouncement
Inputs
Image
Text
File
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Claude Sonnet 4 - Frontier Hybrid Reasoning for High-Volume

Claude Sonnet 4 is a hybrid reasoning large language model from Anthropic that operates in two documented modes: a standard mode for near-instant responses and an extended thinking mode for deeper reasoning. It is the first Claude 4 generation model made available to free-tier users, positioning it as the practical access point for the Claude 4 capability tier.

TraitDetail
Dual-mode architectureSwitches between standard and extended thinking modes per task
Extended thinking with tool useAlternates between reasoning and external tools (e.g. web search) within a single extended thinking session
Parallel tool executionCalls multiple tools simultaneously rather than sequentially
Enhanced steerabilityDocumented as more precisely responsive to operator and user instructions than its predecessor
Shortcut and loophole reduction65% less likely than Sonnet 3.7 to use shortcuts or loopholes on agentic tasks
Thinking summarizationUses a smaller secondary model to condense long thought chains; raw chains available via opt-in Developer Mode
ASL-2 safety classificationReleased under Anthropic's AI Safety Level 2 Standard, the same tier as its predecessor Sonnet 3.7
Free-user availabilityAvailable to Anthropic's free-plan users, unlike Claude Opus 4 which requires a paid plan

Memory When Given File Access

When developers provide Claude Sonnet 4 with local file access, it can extract and save key facts to memory files, enabling continuity across longer sessions. This capability is shared with Claude Opus 4 in the same generation and is not built into the base model - it requires the developer to supply file access in the application layer.

Benchmark Performance

Independent evaluations · Artificial Analysis

29.8%
Intelligence
37.6%
Coding Index
17.6%
Agentic Index

Accuracy & Capability Details

GPQA - Graduate Science77.7%
Humanity's Last Exam10.7%
SciCode - Scientific Coding40.0%
Instruction Following54.7%
Long Context Reasoning69.7%
τ²-Bench - Agentic Tasks64.6%
TerminalBench - System Control31.1%
Specs
Context window
1Mtokens
Input pricing
$3per 1M tokens
Output pricing
$15per 1M tokens
Cached input
$0.30per 1M tokens

Prices in USD.

Key Capabilities & Ratings

Translation & Lang90%
Advanced Math80%
Scientific Biology80%
General Knowledge80%
Physics Problems80%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderAnthropicAnthropicAnthropic
Release DateMay 22, 2025July 24, 2026June 9, 2026
Knowledge CutoffJan 2025May 2026-
Context & Limits
Context Window1M1M1M
Pricing (per 1M tokens)
Input Pricing
$3
Best Input Pricing
$5$10
Output Pricing
$15
Best Output Pricing
$25$50
Modalities
Inputs
imagetextfile
textimage
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index26.0
63.1
Best Intelligence Index
62.1
Coding Index-
78.0
Best Coding Index
76.5
Agentic Index-
59.2
Best Agentic Index
56.6
Claude Sonnet 4
Claude Opus 5
Claude Fable 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

68%
Claude Sonnet 4
GPQA Benchmark
Score: 68%
Claude Sonnet 4
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5
93%
Claude Fable 5
GPQA Benchmark
Score: 93%
Claude Fable 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

4%
Claude Sonnet 4
Humanity's Last Exam
Score: 4%
Claude Sonnet 4
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5
56%
Claude Fable 5
Humanity's Last Exam
Score: 56%
Claude Fable 5

Long Context Reasoning

Logical reasoning over long context windows.

46%
Claude Sonnet 4
Long Context Reasoning
Score: 46%
Claude Sonnet 4
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5
77%
Claude Fable 5
Long Context Reasoning
Score: 77%
Claude Fable 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Back to all models