Toolbit.aiToolbit.ai

Find, compare, and explore the best AI tools to match your specific tasks and use cases.

Explore

  • AI Search
  • Compare ToolsNew
  • Browse Categories
  • Trending Tools
  • Most Popular
  • New Additions

Resources

  • Updates HubNew
  • AI News
  • ModelsNew
  • Blog Articles
  • NewsletterNew

Company

  • Launch a Tool
  • Advertise with Us
  • Guest Post
  • Contact Us
© 2026 Toolbit.ai. All rights reserved.
Privacy PolicyTerms & ConditionsDisclaimer
There's An AI For That favicon
There's An AI For That•The front page of AI for everyone
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Toolbit.ai
Toolbit.ai
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Sign in
  1. Home
  2. Updates
  3. Models
  4. GPT-5.1 Codex
OpenAI
Released November 13, 2025

GPT-5.1 Codex

GPT-5.1 Codex is OpenAI's agentic-coding variant of GPT-5.1, built for Codex harnesses, available only via the Responses API with apply_patch and shell tooling.

Visit OpenAIAnnouncement
Inputs
Text
Image
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

GPT-5.1 Codex - GPT-5.1 Tuned for Agentic Coding in Codex

GPT-5.1 Codex is OpenAI's coding-focused variant of GPT-5.1, built to run inside Codex or similar agent harnesses rather than as a general chat model. It is offered only through the Responses API, where the underlying snapshot is updated over time instead of staying fixed at one release.

TraitDetail
Defining purposeA version of GPT-5.1 optimized for agentic coding tasks inside Codex or Codex-like harnesses, not general-purpose use.
Deployment modelAvailable only through the Responses API, with the model snapshot regularly updated rather than locked at launch.
Companion toolsBuilt to work with apply_patch, a freeform structured-diff tool for file edits without JSON escaping, and a shell tool for running commands directly.

Released Alongside a Smaller Sibling

OpenAI released GPT-5.1 Codex Mini on the same day, positioned as a more cost-efficient option for the same coding and agentic workloads. The pairing reflects a two-tier approach: one Codex model tuned for capability, one tuned for cost, both built on the same GPT-5.1 base.

Benchmark Performance

Independent evaluations · Artificial Analysis

35.6%
Intelligence

Accuracy & Capability Details

GPQA - Graduate Science86.0%
Humanity's Last Exam25.7%
SciCode - Scientific Coding40.2%
Instruction Following70.0%
Long Context Reasoning69.0%
τ²-Bench - Agentic Tasks83.0%
TerminalBench - System Control34.8%
Specs
Context window
400Ktokens
Input pricing
$1.25per 1M tokens
Output pricing
$10per 1M tokens

Prices in USD.

Key Capabilities & Ratings

Advanced Math100%
Logical Logic100%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderOpenAIAnthropicAnthropic
Release DateNovember 13, 2025July 24, 2026June 9, 2026
Knowledge Cutoff-May 2026-
Context & Limits
Context Window400K
1M
Best Context Window
1M
Best Context Window
Pricing (per 1M tokens)
Input Pricing
$1.25
Best Input Pricing
$5$10
Output Pricing
$10
Best Output Pricing
$25$50
Modalities
Inputs
textimage
textimage
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index35.6
63.1
Best Intelligence Index
62.1
Coding Index-
78.0
Best Coding Index
76.5
Agentic Index-
59.2
Best Agentic Index
56.6
GPT-5.1 Codex
Claude Opus 5
Claude Fable 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

86%
GPT-5.1 Codex
GPQA Benchmark
Score: 86%
GPT-5.1 Codex
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5
93%
Claude Fable 5
GPQA Benchmark
Score: 93%
Claude Fable 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

26%
GPT-5.1 Codex
Humanity's Last Exam
Score: 26%
GPT-5.1 Codex
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5
56%
Claude Fable 5
Humanity's Last Exam
Score: 56%
Claude Fable 5

Long Context Reasoning

Logical reasoning over long context windows.

69%
GPT-5.1 Codex
Long Context Reasoning
Score: 69%
GPT-5.1 Codex
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5
77%
Claude Fable 5
Long Context Reasoning
Score: 77%
Claude Fable 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Back to all models