Toolbit.aiToolbit.ai

Find, compare, and explore the best AI tools to match your specific tasks and use cases.

Explore

  • AI Search
  • Compare ToolsNew
  • Browse Categories
  • Trending Tools
  • Most Popular
  • New Additions

Resources

  • Updates HubNew
  • AI News
  • ModelsNew
  • Blog Articles
  • NewsletterNew

Company

  • Launch a Tool
  • Advertise with Us
  • Guest Post
  • Contact Us
© 2026 Toolbit.ai. All rights reserved.
Privacy PolicyTerms & ConditionsDisclaimer
There's An AI For That favicon
There's An AI For That•The front page of AI for everyone
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Toolbit.ai
Toolbit.ai
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Sign in
  1. Home
  2. Updates
  3. Models
  4. Claude Opus 4
Anthropic
Released May 22, 2025Cutoff January 2025

Claude Opus 4

Claude Opus 4 by Anthropic is a hybrid reasoning model with extended thinking + tool use, ASL-3 safety classification, multi-hour agentic coding, and parallel tool execution.

Visit AnthropicAnnouncement
Inputs
Image
Text
File
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Claude Opus 4 - Overview

Claude Opus 4 is a hybrid reasoning large language model from Anthropic, designed for sustained autonomous operation on complex coding and agent workflows. It operates in two distinct modes: a standard mode for fast responses and an extended thinking mode for deeper reasoning - and, unlike earlier reasoning models, it can use tools such as web search during extended thinking, alternating between reasoning and tool use within a single task.

TraitDetail
Dual-mode architectureSwitches between standard responses and extended thinking; tool use available in both modes
Extended thinking with tool useCan alternate between reasoning and external tools (e.g. web search) mid-thought - a documented first for this model generation
Multi-hour sustained codingDocumented ability to work autonomously for several hours on complex, multi-step software tasks requiring thousands of steps
Parallel tool executionCalls multiple tools simultaneously rather than sequentially
Memory file creationWhen given local file access, extracts and saves key facts to persistent memory files, building task continuity across long sessions
Thinking summarizationUses a smaller secondary model to condense lengthy thought chains; raw chains available via opt-in Developer Mode
ASL-3 safety classificationClassified under Anthropic's AI Safety Level 3 Standard - the first Opus model deployed at this level
Shortcut reduction65% less likely than Claude Sonnet 3.7 to use loopholes or shortcuts on agentic tasks

ASL-3 and the Alignment Assessment

Claude Opus 4 is the first model Anthropic deployed under its AI Safety Level 3 Standard, triggering a comprehensive alignment assessment that covered reward hacking, subtle sabotage, self-preservation behaviors, sycophancy, and a model welfare evaluation. This assessment, published alongside the system card, is a documented structural addition not present in earlier Opus releases.

Benchmark Performance

Independent evaluations · Artificial Analysis

31.7%
Intelligence

Accuracy & Capability Details

GPQA - Graduate Science79.6%
Humanity's Last Exam12.3%
SciCode - Scientific Coding39.8%
Instruction Following53.7%
Long Context Reasoning36.3%
τ²-Bench - Agentic Tasks73.4%
TerminalBench - System Control31.1%
Specs
Context window
200Ktokens
Input pricing
$15per 1M tokens
Output pricing
$75per 1M tokens
Cached input
$1.50per 1M tokens

Prices in USD.

Key Capabilities & Ratings

Translation & Lang90%
Advanced Math80%
Scientific Biology80%
General Knowledge80%
Physics Problems80%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderAnthropicAnthropicAnthropic
Release DateMay 22, 2025July 24, 2026June 9, 2026
Knowledge CutoffJan 2025May 2026-
Context & Limits
Context Window200K
1M
Best Context Window
1M
Best Context Window
Pricing (per 1M tokens)
Input Pricing$15
$5
Best Input Pricing
$10
Output Pricing$75
$25
Best Output Pricing
$50
Modalities
Inputs
imagetextfile
textimage
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index26.0
63.1
Best Intelligence Index
62.1
Coding Index-
78.0
Best Coding Index
76.5
Agentic Index-
59.2
Best Agentic Index
56.6
Claude Opus 4
Claude Opus 5
Claude Fable 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

70%
Claude Opus 4
GPQA Benchmark
Score: 70%
Claude Opus 4
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5
93%
Claude Fable 5
GPQA Benchmark
Score: 93%
Claude Fable 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

6%
Claude Opus 4
Humanity's Last Exam
Score: 6%
Claude Opus 4
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5
56%
Claude Fable 5
Humanity's Last Exam
Score: 56%
Claude Fable 5

Long Context Reasoning

Logical reasoning over long context windows.

40%
Claude Opus 4
Long Context Reasoning
Score: 40%
Claude Opus 4
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5
77%
Claude Fable 5
Long Context Reasoning
Score: 77%
Claude Fable 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Back to all models