SpaceXAI
Released August 12, 2026Cutoff February 2026

Grok 4.6

Grok 4.6 by SpaceXAI: flagship code model with 500K context, configurable reasoning effort, minimal hallucinations, and no realtime access without search

Inputs
Text
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Grok 4.6 - Flagship Code and Agentic Model with Configurable Reasoning

Grok 4.6 is xAI's flagship model for code and general-purpose tasks, defined by three stated pillars: agentic tool calling, minimal hallucinations, and configurable reasoning.

TraitDetail
Configurable reasoningReasoning depth adjustable via reasoning_effort parameter
Context window500,000 tokens
Minimal hallucinationsExplicit design goal stated in model documentation

Post-Training Over Scale

Grok 4.6 reuses the same V9 foundation as its predecessor Grok 4.5, with capability gains delivered through improved supervised fine-tuning and reinforcement learning rather than increased parameter scale. The stated goal is matching or exceeding competing frontier models while preserving the inference speed and token efficiency of Grok 4.5. Supplemental training incorporates SpaceX engineering data, excluding ITAR-restricted material.

Benchmark Performance

Independent evaluations · Artificial Analysis

49.3%
Intelligence
75.9%
Coding Index
53.0%
Agentic Index

Accuracy & Capability Details

GPQA - Graduate Science93.5%
Humanity's Last Exam44.1%
SciCode - Scientific Coding53.0%
Long Context Reasoning81.0%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderSpaceXAIAnthropicOpenAI
Release DateAugust 12, 2026September 1, 2026September 3, 2026
Knowledge CutoffFeb 2026-Apr 2026
Context & Limits
Context Window500K-
1.1M
Best Context Window
Pricing (per 1M tokens)
Input Pricing
$2
Best Input Pricing
$10$10
Output Pricing
$6
Best Output Pricing
$50$50
Modalities
Inputs
text
textimagefile
textimage
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index50.6
56.8
Best Intelligence Index
54.7
Coding Index76.8
81.6
Best Coding Index
76.9
Agentic Index53.6
58.2
Best Agentic Index
51.6
Grok 4.6
Claude Fable 5.1
GPT-6 Astra

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

95%
Grok 4.6
GPQA Benchmark
Score: 95%
Grok 4.6
94%
Claude Fable 5.1
GPQA Benchmark
Score: 94%
Claude Fable 5.1
96%
GPT-6 Astra
GPQA Benchmark
Score: 96%
GPT-6 Astra

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

43%
Grok 4.6
Humanity's Last Exam
Score: 43%
Grok 4.6
59%
Claude Fable 5.1
Humanity's Last Exam
Score: 59%
Claude Fable 5.1
55%
GPT-6 Astra
Humanity's Last Exam
Score: 55%
GPT-6 Astra

Long Context Reasoning

Logical reasoning over long context windows.

80%
Grok 4.6
Long Context Reasoning
Score: 80%
Grok 4.6
85%
Claude Fable 5.1
Long Context Reasoning
Score: 85%
Claude Fable 5.1
81%
GPT-6 Astra
Long Context Reasoning
Score: 81%
GPT-6 Astra

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.