Claude Opus 4.1 by Anthropic is a hybrid reasoning model upgrading Opus 4 with improved multi-file refactoring, agentic search, detail tracking, and ASL-3 safety protections.
Capabilities, design details, and architectural traits
Claude Opus 4.1 is an incremental update to Claude Opus 4, narrowly focused on three documented areas: agentic task performance, real-world coding, and reasoning. It is a hybrid reasoning model, toggling between standard and extended thinking mode (up to 64K tokens), which outputs a visible chain-of-thought for harder problems. It carries the same ASL-3 deployment standard as Opus 4, with voluntary safety evaluations confirming a consistent risk profile rather than a full Responsible Scaling Policy reassessment.
| Trait | Detail |
|---|---|
| Multi-file code refactoring | Documented improvement specifically on multi-file refactoring tasks, with precision in pinpointing corrections without introducing new bugs |
| Agentic search and detail tracking | Explicitly improved for in-depth research and data analysis, particularly around detail tracking across long agentic workflows |
| Hybrid reasoning with extended thinking | Supports both fast standard responses and a visible extended thinking mode; TAU-bench scores use a modified scaffold encouraging written reasoning during multi-turn trajectories |
| SWE-bench scaffold | Uses only two tools - a bash tool and a string-replacement file editor - with no planning tool, unlike the Claude 3.7 Sonnet scaffold |
| ASL-3 with incremental safety scope | Deployed under ASL-3 protections; system card is an addendum to the Claude 4 card rather than a full reassessment, as the model did not meet the "notably more capable" threshold |
Anthropic explicitly recommends replacing all Opus 4 usage with Opus 4.1, at the same price and API footprint. The system card formally classifies the changes as incremental improvements in reasoning quality, instruction-following, and overall performance - not a capability-tier shift.
Independent evaluations · Artificial Analysis
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | Anthropic | Anthropic | Anthropic |
| Release Date | August 5, 2025 | July 24, 2026 | June 9, 2026 |
| Knowledge Cutoff | Jan 2025 | May 2026 | - |
| Context & Limits | |||
| Context Window | 200K | 1M Best Context Window | 1M Best Context Window |
| Pricing (per 1M tokens) | |||
| Input Pricing | $15 | $5 Best Input Pricing | $10 |
| Output Pricing | $75 | $25 Best Output Pricing | $50 |
| Modalities | |||
| Inputs | imagetextfile | textimage | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 28.8 | 63.1 Best Intelligence Index | 62.1 |
| Coding Index | - | 78.0 Best Coding Index | 76.5 |
| Agentic Index | - | 59.2 Best Agentic Index | 56.6 |
Graduate-level reasoning and expert Q&A evaluation.
Extremely difficult logical reasoning and knowledge.
Logical reasoning over long context windows.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.