Anthropic
Released May 28, 2026

Claude Opus 4.8

Claude Opus 4.8 by Anthropic: hybrid reasoning model with adaptive thinking, effort control, and a 1M token context window, built for coding and AI agents.

Inputs
Text
Image
File
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Claude Opus 4.8 - Hybrid Reasoning Model Built for Sustained Agentic Work

Claude Opus 4.8 is Anthropic's most capable publicly available model prior to the Fable launch, officially described as a hybrid reasoning model built for serious coding and AI agents. Its defining design principle is adaptive thinking - the model automatically adjusts how much reasoning it applies based on task complexity, spending more computation on harder problems and responding faster to simpler ones. This is always on and user-controllable via effort settings (including an xhigh tier for maximum computation).

What Distinguishes Opus 4.8 From Its Predecessors

TraitWhat it means for Opus 4.8 specifically
Adaptive thinking, always onAutomatically scales reasoning depth per task; users can also set effort explicitly from low to xhigh
Four times fewer unremarked code flawsOfficially documented: Opus 4.8 is around four times less likely than Opus 4.7 to allow flaws in its own code to pass without comment
Alignment scores matching Mythos PreviewEvaluation-confirmed rates of deceptive behavior and cooperation with misuse are similar to Claude Mythos Preview, Anthropic's most aligned model at launch
Fast mode at 2.5x speedFast mode runs at 2.5 times the speed of fast mode on previous Opus versions
1M token context windowSupports a 1 million token context window for sustained, long-running sessions
Fallback target for Fable 5 safety classifiersWhen Claude Fable 5 declines a flagged request, the API automatically reroutes to Opus 4.8 - a documented platform role unique to this model

Honesty as a Documented Behavioral Shift

Anthropically's official announcement frames honesty as the most prominent improvement in Opus 4.8 over 4.7. The model is explicitly trained to flag uncertainties rather than assert unsupported progress - particularly relevant in agentic coding sessions where overconfident claims about task completion are a documented failure mode in prior models.

Benchmark Performance

Independent evaluations · Artificial Analysis

41.8%
Intelligence
74.3%
Coding Index
41.9%
Agentic Index

Accuracy & Capability Details

GPQA - Graduate Science92.0%
Humanity's Last Exam48.7%
SciCode - Scientific Coding54.4%
Instruction Following62.2%
Long Context Reasoning77.7%
τ²-Bench - Agentic Tasks94.4%
TerminalBench - System Control58.3%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderAnthropicAnthropicAnthropic
Release DateMay 28, 2026September 22, 2026September 28, 2026
Knowledge Cutoff--Jun 2026
Context & Limits
Context Window1M1M1M
Pricing (per 1M tokens)
Input Pricing$5$4
$2
Best Input Pricing
Output Pricing$25$20
$10
Best Output Pricing
Modalities
Inputs
textimagefile
textimagefile
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index41.8
57.6
Best Intelligence Index
56.0
Coding Index74.3--
Agentic Index41.9--
Claude Opus 4.8
Claude Opus 5.5
Claude Sonnet 5.5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

49%
Claude Opus 4.8
Humanity's Last Exam
Score: 49%
Claude Opus 4.8
61%
Claude Opus 5.5
Humanity's Last Exam
Score: 61%
Claude Opus 5.5
55%
Claude Sonnet 5.5
Humanity's Last Exam
Score: 55%
Claude Sonnet 5.5

Long Context Reasoning

Logical reasoning over long context windows.

78%
Claude Opus 4.8
Long Context Reasoning
Score: 78%
Claude Opus 4.8
85%
Claude Opus 5.5
Long Context Reasoning
Score: 85%
Claude Opus 5.5
83%
Claude Sonnet 5.5
Long Context Reasoning
Score: 83%
Claude Sonnet 5.5

SciCode Benchmark

Scientific coding and mathematical modeling.

54%
Claude Opus 4.8
SciCode Benchmark
Score: 54%
Claude Opus 4.8
67%
Claude Opus 5.5
SciCode Benchmark
Score: 67%
Claude Opus 5.5
61%
Claude Sonnet 5.5
SciCode Benchmark
Score: 61%
Claude Sonnet 5.5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Explore more from Anthropic

Other models by Anthropic

Top AI Models

Leading alternatives by intelligence score

View all