Toolbit.aiToolbit.ai

Find, compare, and explore the best AI tools to match your specific tasks and use cases.

Explore

  • AI Search
  • Compare ToolsNew
  • Browse Categories
  • Trending Tools
  • Most Popular
  • New Additions

Resources

  • Updates HubNew
  • AI News
  • ModelsNew
  • Blog Articles
  • NewsletterNew

Company

  • Launch a Tool
  • Advertise with Us
  • Guest Post
  • Contact Us
© 2026 Toolbit.ai. All rights reserved.
Privacy PolicyTerms & ConditionsDisclaimer
There's An AI For That favicon
There's An AI For That•The front page of AI for everyone
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Toolbit.ai
Toolbit.ai
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Sign in
  1. Home
  2. Updates
  3. Models
  4. Claude Opus 4.6
Anthropic
Released February 5, 2026

Claude Opus 4.6

Claude Opus 4.6 by Anthropic: first Opus with 1M token context, adaptive thinking, effort controls, context compaction, and lead performance on agentic search and knowledge work.

Visit AnthropicAnnouncement
Inputs
Text
Image
File
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Claude Opus 4.6 - Overview

Claude Opus 4.6 is the Opus generation that introduced three capabilities new to the Opus class simultaneously: adaptive thinking, effort controls, and a 1M token context window in beta. These were not incremental updates to existing features - they each debuted with Opus 4.6. The model is designed around sustained, long-horizon agentic work, with documented improvements in planning, tool use, self-correction during code review, and reliability across large codebases.

TraitWhat it means for Opus 4.6 specifically
Adaptive thinking - first Opus deploymentthinking: {type: "adaptive"} was introduced with Opus 4.6; prior Opus models only had binary extended thinking on/off
Effort controls - four levelslow, medium, high (default), and max effort levels debuted here; Anthropic explicitly documents that high may overthink simple tasks and recommends dialing down via /effort
1M token context - first OpusBeta 1M token context window arrived with Opus 4.6; previous Opus models topped at 200k
Context compaction - first general availabilityAutomatic context summarization near window limits launched alongside Opus 4.6 on the API
ASL-3 deployment with published Sabotage Risk ReportOpus 4.6 is the first model for which Anthropic committed to and published a formal Sabotage Risk Report under its Responsible Scaling Policy, separate from the system card
Agent teams in Claude Code - debutThe ability to assemble multi-agent teams within Claude Code launched with Opus 4.6

The ASL-3 Threshold and Sabotage Risk Report

The Opus 4.6 system card documents that the model roughly reached pre-defined ASL-4 rule-out thresholds on benchmark tasks. Because a clean rule-out was no longer straightforward, Anthropic supplemented the system card with a dedicated Sabotage Risk Report - a formal safety artifact now required for all future models exceeding Opus 4.5's capability level under the Responsible Scaling Policy.

Benchmark Performance

Independent evaluations · Artificial Analysis

44.9%
Intelligence

Accuracy & Capability Details

GPQA - Graduate Science89.6%
Humanity's Last Exam39.9%
SciCode - Scientific Coding51.9%
Instruction Following53.1%
Long Context Reasoning74.3%
τ²-Bench - Agentic Tasks92.1%
TerminalBench - System Control46.2%
Specs
Context window
1Mtokens
Input pricing
$5per 1M tokens
Output pricing
$25per 1M tokens
Cached input
$0.50per 1M tokens

Prices in USD.

Key Capabilities & Ratings

Communication100%
Search Integration90%
Scientific Biology90%
Physics Problems90%
Translation & Lang90%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderAnthropicAnthropicAnthropic
Release DateFebruary 5, 2026July 24, 2026June 9, 2026
Knowledge Cutoff-May 2026-
Context & Limits
Context Window1M1M1M
Pricing (per 1M tokens)
Input Pricing
$5
Best Input Pricing
$5
Best Input Pricing
$10
Output Pricing
$25
Best Output Pricing
$25
Best Output Pricing
$50
Modalities
Inputs
textimagefile
textimage
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index44.9
63.1
Best Intelligence Index
62.1
Coding Index-
78.0
Best Coding Index
76.5
Agentic Index-
59.2
Best Agentic Index
56.6
Claude Opus 4.6
Claude Opus 5
Claude Fable 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

90%
Claude Opus 4.6
GPQA Benchmark
Score: 90%
Claude Opus 4.6
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5
93%
Claude Fable 5
GPQA Benchmark
Score: 93%
Claude Fable 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

40%
Claude Opus 4.6
Humanity's Last Exam
Score: 40%
Claude Opus 4.6
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5
56%
Claude Fable 5
Humanity's Last Exam
Score: 56%
Claude Fable 5

Long Context Reasoning

Logical reasoning over long context windows.

74%
Claude Opus 4.6
Long Context Reasoning
Score: 74%
Claude Opus 4.6
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5
77%
Claude Fable 5
Long Context Reasoning
Score: 77%
Claude Fable 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Back to all models