Toolbit.aiToolbit.ai

Find, compare, and explore the best AI tools to match your specific tasks and use cases.

Explore

  • AI Search
  • Compare ToolsNew
  • Browse Categories
  • Trending Tools
  • Most Popular
  • New Additions

Resources

  • Updates HubNew
  • AI News
  • ModelsNew
  • Blog Articles
  • NewsletterNew

Company

  • Launch a Tool
  • Advertise with Us
  • Guest Post
  • Contact Us
© 2026 Toolbit.ai. All rights reserved.
Privacy PolicyTerms & ConditionsDisclaimer
There's An AI For That favicon
There's An AI For That•The front page of AI for everyone
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Toolbit.ai
Toolbit.ai
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Sign in
  1. Home
  2. Updates
  3. Models
  4. Claude Opus 4.1
Anthropic
Released August 5, 2025Cutoff January 2025

Claude Opus 4.1

Claude Opus 4.1 by Anthropic is a hybrid reasoning model upgrading Opus 4 with improved multi-file refactoring, agentic search, detail tracking, and ASL-3 safety protections.

Visit AnthropicAnnouncement
Inputs
Image
Text
File
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Claude Opus 4.1 - Targeted Precision Upgrade for Agentic Coding and Deep Research

Claude Opus 4.1 is an incremental update to Claude Opus 4, narrowly focused on three documented areas: agentic task performance, real-world coding, and reasoning. It is a hybrid reasoning model, toggling between standard and extended thinking mode (up to 64K tokens), which outputs a visible chain-of-thought for harder problems. It carries the same ASL-3 deployment standard as Opus 4, with voluntary safety evaluations confirming a consistent risk profile rather than a full Responsible Scaling Policy reassessment.

TraitDetail
Multi-file code refactoringDocumented improvement specifically on multi-file refactoring tasks, with precision in pinpointing corrections without introducing new bugs
Agentic search and detail trackingExplicitly improved for in-depth research and data analysis, particularly around detail tracking across long agentic workflows
Hybrid reasoning with extended thinkingSupports both fast standard responses and a visible extended thinking mode; TAU-bench scores use a modified scaffold encouraging written reasoning during multi-turn trajectories
SWE-bench scaffoldUses only two tools - a bash tool and a string-replacement file editor - with no planning tool, unlike the Claude 3.7 Sonnet scaffold
ASL-3 with incremental safety scopeDeployed under ASL-3 protections; system card is an addendum to the Claude 4 card rather than a full reassessment, as the model did not meet the "notably more capable" threshold

Upgrade Path Positioned Over Opus 4

Anthropic explicitly recommends replacing all Opus 4 usage with Opus 4.1, at the same price and API footprint. The system card formally classifies the changes as incremental improvements in reasoning quality, instruction-following, and overall performance - not a capability-tier shift.

Benchmark Performance

Independent evaluations · Artificial Analysis

34.5%
Intelligence

Accuracy & Capability Details

GPQA - Graduate Science80.9%
Humanity's Last Exam12.5%
SciCode - Scientific Coding40.9%
Instruction Following55.4%
Long Context Reasoning73.3%
τ²-Bench - Agentic Tasks71.4%
TerminalBench - System Control34.3%
Specs
Context window
200Ktokens
Input pricing
$15per 1M tokens
Output pricing
$75per 1M tokens
Cached input
$1.50per 1M tokens

Prices in USD.

Key Capabilities & Ratings

Translation & Lang90%
Advanced Math80%
Visual Reasoning80%
Scientific Biology80%
General Knowledge80%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderAnthropicAnthropicAnthropic
Release DateAugust 5, 2025July 24, 2026June 9, 2026
Knowledge CutoffJan 2025May 2026-
Context & Limits
Context Window200K
1M
Best Context Window
1M
Best Context Window
Pricing (per 1M tokens)
Input Pricing$15
$5
Best Input Pricing
$10
Output Pricing$75
$25
Best Output Pricing
$50
Modalities
Inputs
imagetextfile
textimage
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index28.8
63.1
Best Intelligence Index
62.1
Coding Index-
78.0
Best Coding Index
76.5
Agentic Index-
59.2
Best Agentic Index
56.6
Claude Opus 5
Claude Fable 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

N/A
Claude Opus 4.1
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5
93%
Claude Fable 5
GPQA Benchmark
Score: 93%
Claude Fable 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

N/A
Claude Opus 4.1
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5
56%
Claude Fable 5
Humanity's Last Exam
Score: 56%
Claude Fable 5

Long Context Reasoning

Logical reasoning over long context windows.

N/A
Claude Opus 4.1
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5
77%
Claude Fable 5
Long Context Reasoning
Score: 77%
Claude Fable 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Back to all models