Toolbit.aiToolbit.ai

Find, compare, and explore the best AI tools to match your specific tasks and use cases.

Explore

  • AI Search
  • Compare ToolsNew
  • Browse Categories
  • Trending Tools
  • Most Popular
  • New Additions

Resources

  • Updates HubNew
  • AI News
  • ModelsNew
  • Blog Articles
  • NewsletterNew

Company

  • Launch a Tool
  • Advertise with Us
  • Guest Post
  • Contact Us
© 2026 Toolbit.ai. All rights reserved.
Privacy PolicyTerms & ConditionsDisclaimer
There's An AI For That favicon
There's An AI For That•The front page of AI for everyone
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Toolbit.ai
Toolbit.ai
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Sign in
  1. Home
  2. Updates
  3. Models
  4. Claude Sonnet 4.5
Anthropic
Released September 29, 2025Cutoff January 2025

Claude Sonnet 4.5

Claude Sonnet 4.5 by Anthropic is a hybrid reasoning model built for agentic coding, computer use, and long-running tasks with extended thinking and ASL-3 safety protections.

Visit AnthropicAnnouncement
Inputs
Text
Image
File
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Claude Sonnet 4.5 - Anthropic's Frontier for Agentic Coding

Claude Sonnet 4.5 is a hybrid reasoning model from Anthropic, designed as the primary workhorse for agentic tasks, software engineering, and direct computer use. It operates in two modes: a fast default response mode and extended thinking mode, where it outputs a visible chain-of-thought for harder problems. It is the first Claude model deployed under AI Safety Level 3 (ASL-3) protections, with classifiers specifically targeting CBRN-related inputs and outputs.

TraitDetail
Dual operating modesToggles between standard response and extended thinking, with a visible reasoning chain for complex problems
30+ hour task enduranceDocumented ability to sustain focus on multi-step, long-horizon tasks across agentic loops
Computer use leadershipLeads on OSWorld at 61.4%, designed to navigate browsers, click buttons, fill forms, and recover from errors
ASL-3 deploymentReleased under Anthropic's AI Safety Level 3 standard with CBRN classifiers and prompt-injection defenses for agentic use
Most aligned frontier model (at release)System card includes first use of mechanistic interpretability to evaluate alignment, with documented reductions in sycophancy, deception, and power-seeking
Context editing and memory toolSupports extended autonomous agent sessions via an API-level memory tool and context editing feature

Built Around the Claude Agent SDK

Sonnet 4.5 launched alongside the Claude Agent SDK, giving developers the same agent-building infrastructure Anthropic uses internally for Claude Code. The model is designed to be orchestrated as well as to orchestrate - Anthropic explicitly documented a pattern where Sonnet 4.5 breaks down problems and coordinates multiple Haiku 4.5 instances as sub-agents in parallel.

Benchmark Performance

Independent evaluations · Artificial Analysis

37.4%
Intelligence
52.1%
Coding Index
26.4%
Agentic Index

Accuracy & Capability Details

GPQA - Graduate Science83.4%
Humanity's Last Exam17.8%
SciCode - Scientific Coding44.7%
Instruction Following57.3%
Long Context Reasoning68.3%
τ²-Bench - Agentic Tasks78.1%
TerminalBench - System Control35.6%
Specs
Context window
1Mtokens
Input pricing
$3per 1M tokens
Output pricing
$15per 1M tokens
Cached input
$0.30per 1M tokens

Prices in USD.

Key Capabilities & Ratings

Advanced Math90%
Translation & Lang90%
Scientific Biology80%
General Knowledge80%
Physics Problems80%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderAnthropicAnthropicAnthropic
Release DateSeptember 29, 2025July 24, 2026June 9, 2026
Knowledge CutoffJan 2025May 2026-
Context & Limits
Context Window1M1M1M
Pricing (per 1M tokens)
Input Pricing
$3
Best Input Pricing
$5$10
Output Pricing
$15
Best Output Pricing
$25$50
Modalities
Inputs
textimagefile
textimage
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index29.9
63.1
Best Intelligence Index
62.1
Coding Index-
78.0
Best Coding Index
76.5
Agentic Index-
59.2
Best Agentic Index
56.6
Claude Sonnet 4.5
Claude Opus 5
Claude Fable 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

73%
Claude Sonnet 4.5
GPQA Benchmark
Score: 73%
Claude Sonnet 4.5
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5
93%
Claude Fable 5
GPQA Benchmark
Score: 93%
Claude Fable 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

7%
Claude Sonnet 4.5
Humanity's Last Exam
Score: 7%
Claude Sonnet 4.5
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5
56%
Claude Fable 5
Humanity's Last Exam
Score: 56%
Claude Fable 5

Long Context Reasoning

Logical reasoning over long context windows.

52%
Claude Sonnet 4.5
Long Context Reasoning
Score: 52%
Claude Sonnet 4.5
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5
77%
Claude Fable 5
Long Context Reasoning
Score: 77%
Claude Fable 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Back to all models