Grok 4.20 by xAI delivers industry-leading speed, agentic tool calling, strict prompt adherence, and the lowest documented hallucination rate across a 2M token context window.
Capabilities, design details, and architectural traits
Grok 4.20 is xAI's high-performance model positioned around agentic tool calling combined with a documented emphasis on minimizing hallucinations and strict prompt adherence. It is built to deliver consistently precise, truthful responses at speed.
The 2M token context window is the largest officially documented for any Grok model, enabling ingestion of long documents or extended conversation histories in a single request without chunking.
Officially positioned as a model that pairs throughput and context depth with agentic tool use — suited for workflows where the model must call external tools rapidly while maintaining output accuracy.
Independent evaluations · Artificial Analysis
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | xAI | Anthropic | Anthropic |
| Release Date | March 10, 2026 | July 24, 2026 | June 9, 2026 |
| Knowledge Cutoff | Sep 2025 | May 2026 | - |
| Context & Limits | |||
| Context Window | 2M Best Context Window | 1M | 1M |
| Pricing (per 1M tokens) | |||
| Input Pricing | $1.25 Best Input Pricing | $5 | $10 |
| Output Pricing | $2.50 Best Output Pricing | $25 | $50 |
| Modalities | |||
| Inputs | textimagefile | textimage | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 38.0 | 63.1 Best Intelligence Index | 62.1 |
| Coding Index | - | 78.0 Best Coding Index | 76.5 |
| Agentic Index | - | 59.2 Best Agentic Index | 56.6 |
Graduate-level reasoning and expert Q&A evaluation.
Extremely difficult logical reasoning and knowledge.
Logical reasoning over long context windows.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.