SpaceXAI
Released August 12, 2026Cutoff February 2026

Grok 4.6

Grok 4.6 by SpaceXAI: flagship code model with 500K context, configurable reasoning effort, minimal hallucinations, and no realtime access without search

Inputs
Text
Image
File
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Grok 4.6 - Flagship Code and Agentic Model with Configurable Reasoning

Grok 4.6 is xAI's flagship model for code and general-purpose tasks, defined by three stated pillars: agentic tool calling, minimal hallucinations, and configurable reasoning.

TraitDetail
Configurable reasoningReasoning depth adjustable via reasoning_effort parameter
Context window500,000 tokens
Minimal hallucinationsExplicit design goal stated in model documentation

Post-Training Over Scale

Grok 4.6 reuses the same V9 foundation as its predecessor Grok 4.5, with capability gains delivered through improved supervised fine-tuning and reinforcement learning rather than increased parameter scale. The stated goal is matching or exceeding competing frontier models while preserving the inference speed and token efficiency of Grok 4.5. Supplemental training incorporates SpaceX engineering data, excluding ITAR-restricted material.

Benchmark Performance

Independent evaluations · Artificial Analysis

44.2%
Intelligence
75.9%
Coding Index
52.5%
Agentic Index

Accuracy & Capability Details

GPQA - Graduate Science93.5%
Humanity's Last Exam44.1%
SciCode - Scientific Coding53.0%
Long Context Reasoning81.0%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderSpaceXAIAnthropicAnthropic
Release DateAugust 12, 2026September 22, 2026September 28, 2026
Knowledge CutoffFeb 2026-Jun 2026
Context & Limits
Context Window500K
1M
Best Context Window
1M
Best Context Window
Pricing (per 1M tokens)
Input Pricing
$2
Best Input Pricing
$4
$2
Best Input Pricing
Output Pricing
$6
Best Output Pricing
$20$10
Modalities
Inputs
textimagefile
textimagefile
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index44.3
57.6
Best Intelligence Index
56.0
Coding Index76.8--
Agentic Index53.0--
Grok 4.6
Claude Opus 5.5
Claude Sonnet 5.5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

43%
Grok 4.6
Humanity's Last Exam
Score: 43%
Grok 4.6
61%
Claude Opus 5.5
Humanity's Last Exam
Score: 61%
Claude Opus 5.5
55%
Claude Sonnet 5.5
Humanity's Last Exam
Score: 55%
Claude Sonnet 5.5

Long Context Reasoning

Logical reasoning over long context windows.

80%
Grok 4.6
Long Context Reasoning
Score: 80%
Grok 4.6
85%
Claude Opus 5.5
Long Context Reasoning
Score: 85%
Claude Opus 5.5
83%
Claude Sonnet 5.5
Long Context Reasoning
Score: 83%
Claude Sonnet 5.5

SciCode Benchmark

Scientific coding and mathematical modeling.

56%
Grok 4.6
SciCode Benchmark
Score: 56%
Grok 4.6
67%
Claude Opus 5.5
SciCode Benchmark
Score: 67%
Claude Opus 5.5
61%
Claude Sonnet 5.5
SciCode Benchmark
Score: 61%
Claude Sonnet 5.5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Explore more from SpaceXAI

Other models by SpaceXAI

Top AI Models

Leading alternatives by intelligence score

View all