GPT 5.3-Codex by OpenAI is a Codex-native agent merging frontier coding with GPT-5.2 reasoning, real-time steering, computer use, and 400K-token context.
Capabilities, design details, and architectural traits
GPT 5.3-Codex is OpenAI's most capable agentic coding model, and the first in the GPT-5 family to directly combine frontier coding performance with professional knowledge and general reasoning in a single model. Uniquely, early versions were used by the Codex team to debug their own training run, manage deployment, and diagnose evaluations - making GPT 5.3-Codex the first model documented as having been instrumental in creating itself.
| Trait | Detail |
|---|---|
| Merged training stacks | First model to combine the GPT-5.2-Codex coding stack with the GPT-5.2 reasoning and professional knowledge stack, rather than specializing in one |
| Real-time steering | Users can redirect the agent mid-task without losing context - described as interacting "much like a colleague" while it works |
| Computer use across OSWorld | Extends beyond terminal and code to visual desktop environments; supports OS-level task completion via screenshot-based interaction |
| Full software lifecycle scope | Targets debugging, deploying, monitoring, writing PRDs, editing copy, user research, metrics, slide decks, and spreadsheets - not code generation only |
| High-capability cybersecurity classification | First OpenAI model classified as High capability for cybersecurity under the Preparedness Framework; trained to identify software vulnerabilities |
| Reasoning effort control | Supports low, medium, high, and xhigh effort settings, allowing per-request compute tuning in Codex and API environments |
| Self-referential development loop | Used to optimize its own training harness, root-cause inference bugs, scale GPU clusters during launch, and analyze its own session logs |
GPT 5.3-Codex is designed for the Codex environment - the app, CLI, and IDE extension - and available via both Chat Completions and Responses API. Its 400k token context window and 128k token max output support long-horizon tasks that iterate over millions of tokens, such as autonomously building multi-map games or full financial presentation decks from a single prompt.
Independent evaluations · Artificial Analysis
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | OpenAI | Anthropic | Anthropic |
| Release Date | February 5, 2026 | July 24, 2026 | June 9, 2026 |
| Knowledge Cutoff | - | May 2026 | - |
| Context & Limits | |||
| Context Window | 400K | 1M Best Context Window | 1M Best Context Window |
| Pricing (per 1M tokens) | |||
| Input Pricing | $1.75 Best Input Pricing | $5 | $10 |
| Output Pricing | $14 Best Output Pricing | $25 | $50 |
| Modalities | |||
| Inputs | textimagefile | textimage | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 45.5 | 63.1 Best Intelligence Index | 62.1 |
| Coding Index | - | 78.0 Best Coding Index | 76.5 |
| Agentic Index | - | 59.2 Best Agentic Index | 56.6 |
Graduate-level reasoning and expert Q&A evaluation.
Extremely difficult logical reasoning and knowledge.
Logical reasoning over long context windows.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.