OpenAI's GPT-4: a large multimodal model accepting image and text inputs. Features predictable scaling, OpenAI Evals, and advanced reasoning.
Capabilities, design details, and architectural traits
GPT 4 is described by OpenAI as a large multimodal model that accepts image and text input and produces text output only. It marks a specific point in OpenAI's own development process, where training behavior became something the company could predict ahead of time.
| Trait | Detail |
|---|---|
| Input and output split | Accepts both image and text as input, but generates text output only, unlike later models built for output across multiple modalities. |
| Predictable training run | Documented by OpenAI as its first large model whose training performance could be accurately predicted ahead of time, based on lessons from an earlier GPT-3.5 test run. |
| Six month alignment period | OpenAI states it spent six months iteratively aligning GPT 4 using lessons from adversarial testing and ChatGPT, aimed at factuality, steerability, and guardrail adherence. |
| Staged image rollout | At launch, text input shipped broadly through ChatGPT and the API with a waitlist, while image input was prepared for wider availability through a single partner first. |
OpenAI open sourced OpenAI Evals alongside GPT 4, a framework for automated evaluation of AI model performance, built so anyone could report shortcomings and help guide further improvements.
Independent evaluations · Artificial Analysis
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | OpenAI | Anthropic | Anthropic |
| Release Date | March 14, 2023 | July 24, 2026 | June 9, 2026 |
| Knowledge Cutoff | Sep 2021 | May 2026 | - |
| Context & Limits | |||
| Context Window | 8K | 1M Best Context Window | 1M Best Context Window |
| Pricing (per 1M tokens) | |||
| Input Pricing | $30 | $5 Best Input Pricing | $10 |
| Output Pricing | $60 | $25 Best Output Pricing | $50 |
| Modalities | |||
| Inputs | text | textimage | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 6.8 | 63.1 Best Intelligence Index | 62.1 |
| Coding Index | 13.1 | 78.0 Best Coding Index | 76.5 |
| Agentic Index | - | 59.2 Best Agentic Index | 56.6 |
Graduate-level reasoning and expert Q&A evaluation.
Extremely difficult logical reasoning and knowledge.
Evaluation of strict instruction following.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.