Toolbit.aiToolbit.ai

Find, compare, and explore the best AI tools to match your specific tasks and use cases.

Explore

  • AI Search
  • Compare ToolsNew
  • Browse Categories
  • Trending Tools
  • Most Popular
  • New Additions

Resources

  • Updates HubNew
  • AI News
  • ModelsNew
  • Blog Articles
  • NewsletterNew

Company

  • Launch a Tool
  • Advertise with Us
  • Guest Post
  • Contact Us
© 2026 Toolbit.ai. All rights reserved.
Privacy PolicyTerms & ConditionsDisclaimer
There's An AI For That favicon
There's An AI For That•The front page of AI for everyone
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Toolbit.ai
Toolbit.ai
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Sign in
  1. Home
  2. Updates
  3. Models
  4. Grok 4.5
SpaceXAI
Released July 8, 2026

Grok 4.5

Grok 4.5 by SpaceXAI is a frontier model for coding, agentic tasks, and knowledge work with configurable reasoning, 500k context, and built-in tools.

Visit SpaceXAIAnnouncement
Inputs
Text
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Grok 4.5 - Frontier model for coding, agentic tasks, and knowledge work

Grok 4.5 is SpaceXAI's flagship model, positioned as the most intelligent and fastest model the company has built. It was trained in SpaceXAI's Memphis data centers using new datasets spanning science, engineering, and math, and is designed as a workhorse for coding, agentic tool calling, and routine knowledge work.

TraitDetail
Configurable reasoningLow, medium, or high effort selectable via reasoning_effort, defaulting to high
500k token context windowSupports long conversations and extended agent loops with context compaction
Built-in tool suiteFunction calling, web search, X search, and code execution available as server-side tools
Minimal hallucinationsExplicitly listed as a defining model characteristic alongside agentic tool calling
Grok Build defaultServes as the default model for SpaceXAI's coding agent on both API and CLI
Office add-in integrationDefault model in Word, PowerPoint, and Excel add-ins
Twice greater token efficiencyClaimed relative to other leading models, reducing per-task token cost

Reasoning and tool integration

Grok 4.5 exposes reasoning_effort as a first-class parameter, letting developers trade latency for depth on a per-request basis. Its tool ecosystem is tightly coupled: web search and X search provide realtime data beyond the knowledge cutoff, while code execution enables in-context programmatic reasoning for coding and math tasks.

Deployment surface

Beyond the xAI API, Grok 4.5 ships as the default model in Grok Build (SpaceXAI's coding agent), in Cursor on all plans, and in Office add-ins. It is also accessible through model gateways including OpenRouter, Vercel, Cloudflare, Snowflake, and Databricks Mosaic.

Benchmark Performance

Independent evaluations · Artificial Analysis

55.8%
Intelligence
72.4%
Coding Index
48.9%
Agentic Index

Accuracy & Capability Details

GPQA - Graduate Science93.1%
Humanity's Last Exam42.7%
SciCode - Scientific Coding54.1%
Long Context Reasoning74.0%
Specs
Context window
-tokens
Input pricing
$2per 1M tokens
Output pricing
$6per 1M tokens
Cached input
$0.30per 1M tokens

Prices in USD.

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderSpaceXAIAnthropicAnthropic
Release DateJuly 8, 2026July 24, 2026June 9, 2026
Knowledge Cutoff-May 2026-
Context & Limits
Context Window-1M1M
Pricing (per 1M tokens)
Input Pricing
$2
Best Input Pricing
$5$10
Output Pricing
$6
Best Output Pricing
$25$50
Modalities
Inputs
text
textimage
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index55.8
63.1
Best Intelligence Index
62.1
Coding Index72.4
78.0
Best Coding Index
76.5
Agentic Index48.9
59.2
Best Agentic Index
56.6
Grok 4.5
Claude Opus 5
Claude Fable 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

93%
Grok 4.5
GPQA Benchmark
Score: 93%
Grok 4.5
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5
93%
Claude Fable 5
GPQA Benchmark
Score: 93%
Claude Fable 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

43%
Grok 4.5
Humanity's Last Exam
Score: 43%
Grok 4.5
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5
56%
Claude Fable 5
Humanity's Last Exam
Score: 56%
Claude Fable 5

Long Context Reasoning

Logical reasoning over long context windows.

74%
Grok 4.5
Long Context Reasoning
Score: 74%
Grok 4.5
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5
77%
Claude Fable 5
Long Context Reasoning
Score: 77%
Claude Fable 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Back to all models