Toolbit.aiToolbit.ai

Find, compare, and explore the best AI tools to match your specific tasks and use cases.

Explore

  • AI Search
  • Compare ToolsNew
  • Browse Categories
  • Trending Tools
  • Most Popular
  • New Additions

Resources

  • Updates HubNew
  • AI News
  • ModelsNew
  • Blog Articles
  • NewsletterNew

Company

  • Launch a Tool
  • Advertise with Us
  • Guest Post
  • Contact Us
© 2026 Toolbit.ai. All rights reserved.
Privacy PolicyTerms & ConditionsDisclaimer
There's An AI For That favicon
There's An AI For That•The front page of AI for everyone
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Toolbit.ai
Toolbit.ai
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Sign in
  1. Home
  2. Updates
  3. Models
  4. GPT-5.1 Codex Mini
OpenAI
Released November 13, 2025

GPT-5.1 Codex Mini

GPT-5.1 Codex Mini is OpenAI's smaller, faster version of GPT-5.1 Codex, callable only via the Responses API for agentic coding in Codex harnesses.

Visit OpenAIAnnouncement
Inputs
Image
Text
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

GPT-5.1 Codex Mini - The Smaller, Faster Version of GPT-5.1 Codex

GPT-5.1 Codex Mini is documented as a smaller, faster version of GPT-5.1 Codex, OpenAI's coding-focused variant of GPT-5.1. It carries the same purpose as its larger sibling, agentic coding inside Codex or Codex-like harnesses, but in a lighter form.

TraitDetail
Defining purposeA smaller, faster version of GPT-5.1 Codex, built for agentic coding tasks in Codex or Codex-like harnesses rather than general-purpose use.
Deployment modelCallable only through the Responses API, the same access restriction OpenAI documents for its other Codex-tuned models.
Release pairingReleased together with GPT-5.1 and GPT-5.1 Codex as the lighter, lower-cost option in that same model set.

A Lighter Tier in the Codex Lineup

OpenAI's developer announcement groups gpt-5.1-codex and gpt-5.1-codex-mini together as the same family of models optimized for long-running, agentic coding work. The Mini variant exists specifically to offer that same agentic-coding tuning in a smaller, faster form rather than as a separate use case.

Benchmark Performance

Independent evaluations · Artificial Analysis

31.3%
Intelligence

Accuracy & Capability Details

GPQA - Graduate Science81.3%
Humanity's Last Exam18.5%
SciCode - Scientific Coding42.6%
Instruction Following67.9%
Long Context Reasoning65.0%
τ²-Bench - Agentic Tasks62.9%
TerminalBench - System Control33.3%
Specs
Context window
400Ktokens
Input pricing
$0.25per 1M tokens
Output pricing
$2per 1M tokens

Prices in USD.

Key Capabilities & Ratings

Advanced Math40%
Logical Logic40%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderOpenAIAnthropicAnthropic
Release DateNovember 13, 2025July 24, 2026June 9, 2026
Knowledge Cutoff-May 2026-
Context & Limits
Context Window400K
1M
Best Context Window
1M
Best Context Window
Pricing (per 1M tokens)
Input Pricing
$0.25
Best Input Pricing
$5$10
Output Pricing
$2
Best Output Pricing
$25$50
Modalities
Inputs
imagetext
textimage
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index31.3
63.1
Best Intelligence Index
62.1
Coding Index-
78.0
Best Coding Index
76.5
Agentic Index-
59.2
Best Agentic Index
56.6
GPT-5.1 Codex Mini
Claude Opus 5
Claude Fable 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

81%
GPT-5.1 Codex Mini
GPQA Benchmark
Score: 81%
GPT-5.1 Codex Mini
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5
93%
Claude Fable 5
GPQA Benchmark
Score: 93%
Claude Fable 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

19%
GPT-5.1 Codex Mini
Humanity's Last Exam
Score: 19%
GPT-5.1 Codex Mini
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5
56%
Claude Fable 5
Humanity's Last Exam
Score: 56%
Claude Fable 5

Long Context Reasoning

Logical reasoning over long context windows.

65%
GPT-5.1 Codex Mini
Long Context Reasoning
Score: 65%
GPT-5.1 Codex Mini
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5
77%
Claude Fable 5
Long Context Reasoning
Score: 77%
Claude Fable 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Back to all models