Claude Opus 4 by Anthropic is a hybrid reasoning model with extended thinking + tool use, ASL-3 safety classification, multi-hour agentic coding, and parallel tool execution.
Capabilities, design details, and architectural traits
Claude Opus 4 is a hybrid reasoning large language model from Anthropic, designed for sustained autonomous operation on complex coding and agent workflows. It operates in two distinct modes: a standard mode for fast responses and an extended thinking mode for deeper reasoning - and, unlike earlier reasoning models, it can use tools such as web search during extended thinking, alternating between reasoning and tool use within a single task.
| Trait | Detail |
|---|---|
| Dual-mode architecture | Switches between standard responses and extended thinking; tool use available in both modes |
| Extended thinking with tool use | Can alternate between reasoning and external tools (e.g. web search) mid-thought - a documented first for this model generation |
| Multi-hour sustained coding | Documented ability to work autonomously for several hours on complex, multi-step software tasks requiring thousands of steps |
| Parallel tool execution | Calls multiple tools simultaneously rather than sequentially |
| Memory file creation | When given local file access, extracts and saves key facts to persistent memory files, building task continuity across long sessions |
| Thinking summarization | Uses a smaller secondary model to condense lengthy thought chains; raw chains available via opt-in Developer Mode |
| ASL-3 safety classification | Classified under Anthropic's AI Safety Level 3 Standard - the first Opus model deployed at this level |
| Shortcut reduction | 65% less likely than Claude Sonnet 3.7 to use loopholes or shortcuts on agentic tasks |
Claude Opus 4 is the first model Anthropic deployed under its AI Safety Level 3 Standard, triggering a comprehensive alignment assessment that covered reward hacking, subtle sabotage, self-preservation behaviors, sycophancy, and a model welfare evaluation. This assessment, published alongside the system card, is a documented structural addition not present in earlier Opus releases.
Independent evaluations · Artificial Analysis
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | Anthropic | Anthropic | Anthropic |
| Release Date | May 22, 2025 | July 24, 2026 | June 9, 2026 |
| Knowledge Cutoff | Jan 2025 | May 2026 | - |
| Context & Limits | |||
| Context Window | 200K | 1M Best Context Window | 1M Best Context Window |
| Pricing (per 1M tokens) | |||
| Input Pricing | $15 | $5 Best Input Pricing | $10 |
| Output Pricing | $75 | $25 Best Output Pricing | $50 |
| Modalities | |||
| Inputs | imagetextfile | textimage | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 26.0 | 63.1 Best Intelligence Index | 62.1 |
| Coding Index | - | 78.0 Best Coding Index | 76.5 |
| Agentic Index | - | 59.2 Best Agentic Index | 56.6 |
Graduate-level reasoning and expert Q&A evaluation.
Extremely difficult logical reasoning and knowledge.
Logical reasoning over long context windows.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.