GPT 5.4 Pro by OpenAI delivers maximum-compute reasoning, native computer use, 1M-token context, tool search, and 89.3% BrowseComp for complex professional tasks.
Capabilities, design details, and architectural traits
GPT 5.4 Pro is OpenAI's highest-performance variant of GPT 5.4, designed to spend more compute per request to produce smarter and more precise responses on demanding professional tasks. It is available exclusively through the Responses API, enabling multi-turn model interactions before the final reply is delivered.
| Trait | Detail |
|---|---|
| Compute scaling | Uses more compute per request to think harder; reasoning.effort supports medium (default), high, and xhigh |
| Responses API only | Exclusively available via the Responses API - not the Chat Completions endpoint - to support multi-turn internal reasoning before output |
| Background mode | Designed for long-running requests; background mode recommended to prevent timeouts on the most complex tasks |
| Tool search | Dynamically retrieves tool definitions on demand rather than loading all tool context upfront, reducing token overhead by up to 47% on large MCP ecosystems |
| Native computer use | Directly issues mouse and keyboard commands from screenshots; behavior is steerable via developer messages and configurable confirmation policies |
GPT 5.4 Pro is built around the idea that professional-grade work requires iterative, tool-heavy execution over long contexts. Its exclusive Responses API access supports the infrastructure for multi-turn reasoning loops, and its tool search capability lets agents operate across large ecosystems - such as MCP servers with tens of thousands of tool tokens - without bloating the context on every request.
GPT 5.4 Pro targets structured professional output: legal analysis, financial modeling, spreadsheet creation, and multi-step research. Its CoT monitorability is explicitly low - meaning the model cannot deliberately hide its chain-of-thought reasoning - which is documented as a safety property, not a limitation.
No independent benchmark data is available for this model yet.
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | OpenAI | Anthropic | Anthropic |
| Release Date | March 5, 2026 | July 24, 2026 | June 9, 2026 |
| Knowledge Cutoff | - | May 2026 | - |
| Context & Limits | |||
| Context Window | 1.1M Best Context Window | 1M | 1M |
| Pricing (per 1M tokens) | |||
| Input Pricing | $30 | $5 Best Input Pricing | $10 |
| Output Pricing | $180 | $25 Best Output Pricing | $50 |
| Modalities | |||
| Inputs | textimagefile | textimage | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | - | 63.1 Best Intelligence Index | 62.1 |
| Coding Index | - | 78.0 Best Coding Index | 76.5 |
| Agentic Index | - | 59.2 Best Agentic Index | 56.6 |
Graduate-level reasoning and expert Q&A evaluation.
Extremely difficult logical reasoning and knowledge.
Logical reasoning over long context windows.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.