gpt-oss-120b is OpenAI's open-weight reasoning model under Apache 2.0, built for production-scale agentic use with adjustable reasoning and harmony format.
Capabilities, design details, and architectural traits
gpt-oss-120b is an open-weight reasoning model released under the Apache 2.0 license. OpenAI positions it as the larger of two gpt-oss models, intended for production, general purpose, high reasoning work rather than lightweight or local deployment.
| Trait | Detail |
|---|---|
| Defining purpose | Positioned by OpenAI for production, general purpose, high reasoning use cases, as the larger of the two gpt-oss models. |
| Format requirement | Trained on OpenAI's harmony response format; OpenAI states it should only be used in that format, since it will not work correctly otherwise. |
| Adversarial safety testing | OpenAI specifically adversarially fine-tuned gpt-oss-120b under its Preparedness Framework to test whether determined attackers could push it to High capability in biological, chemical, or cyber risk before release. |
| Agentic tool use | Supports adjustable reasoning effort across low, medium, and high settings, and is built for agentic workflows with tool use such as web search and Python code execution. |
OpenAI states that open-weight models like gpt-oss-120b carry a different risk profile than its proprietary models. Once released, a determined attacker could fine-tune the weights to bypass safety refusals, and OpenAI would have no way to add further mitigations or revoke access after the fact.
Independent evaluations · Artificial Analysis
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | OpenAI | Anthropic | Anthropic |
| Release Date | August 5, 2025 | July 24, 2026 | June 9, 2026 |
| Knowledge Cutoff | Jun 2024 | May 2026 | - |
| Context & Limits | |||
| Context Window | 131K | 1M Best Context Window | 1M Best Context Window |
| Pricing (per 1M tokens) | |||
| Input Pricing | $0.15 Best Input Pricing | $5 | $10 |
| Output Pricing | $0.60 Best Output Pricing | $25 | $50 |
| Modalities | |||
| Inputs | text | textimage | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 24.1 | 63.1 Best Intelligence Index | 62.1 |
| Coding Index | 30.4 | 78.0 Best Coding Index | 76.5 |
| Agentic Index | 13.4 | 59.2 Best Agentic Index | 56.6 |
Graduate-level reasoning and expert Q&A evaluation.
Extremely difficult logical reasoning and knowledge.
Logical reasoning over long context windows.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.