Toolbit.aiToolbit.ai

Find, compare, and explore the best AI tools to match your specific tasks and use cases.

Explore

  • AI Search
  • Compare ToolsNew
  • Browse Categories
  • Trending Tools
  • Most Popular
  • New Additions

Resources

  • Updates HubNew
  • AI News
  • ModelsNew
  • Blog Articles
  • NewsletterNew

Company

  • Launch a Tool
  • Advertise with Us
  • Guest Post
  • Contact Us
© 2026 Toolbit.ai. All rights reserved.
Privacy PolicyTerms & ConditionsDisclaimer
There's An AI For That favicon
There's An AI For That•The front page of AI for everyone
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Toolbit.ai
Toolbit.ai
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Sign in
  1. Home
  2. Updates
  3. Models
  4. Gemma 4 31B
Google
Released April 2, 2026

Gemma 4 31B

Google Gemma 4 31B: dense open-weights model built from Gemini 3 research, designed for consumer GPUs and workstations, with built-in thinking, MTP, and fine-tuning support.

Visit GoogleAnnouncement
Inputs
Image
Text
Video
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Gemma 4 31B - Dense Open Model at the Workstation Performance Tier

Gemma 4 31B is Google's dense model in the Gemma 4 family, explicitly positioned to bridge server-grade performance and local execution on consumer GPUs and workstations. Within the Gemma 4 lineup, it is the variant that maximizes raw quality and serves as the primary foundation for fine-tuning, contrasting with the 26B MoE variant which optimizes for inference latency. The entire Gemma 4 family is built from Gemini 3 research and released as open weights.

TraitDetail
ArchitectureDense 31B - not Mixture-of-Experts; all parameters active during inference, prioritizing raw output quality over speed
Deployment targetConsumer GPUs and workstations; explicitly positioned for local-first AI development and fine-tuning
Multi-Token Prediction (MTP)Supported; universally recommended for all tasks on GPU backends for this model size
Built-in thinking modeStep-by-step reasoning before answering, available across the Gemma 4 family
Quantization checkpointsQAT checkpoints available in compressed-tensors (-w4a16-ct) format for optimized cloud serving

Hybrid Attention Across the Gemma 4 Family

All Gemma 4 models, including the 31B, use a hybrid attention mechanism that interleaves local sliding window attention with full global attention. The final layer is always global. Global layers use unified Keys and Values with Proportional RoPE (p-RoPE) for long-context memory optimization.

Benchmark Performance

Independent evaluations · Artificial Analysis

29.7%
Intelligence
43.4%
Coding Index
14.4%
Agentic Index

Accuracy & Capability Details

GPQA - Graduate Science85.7%
Humanity's Last Exam23.6%
SciCode - Scientific Coding43.4%
Instruction Following75.6%
Long Context Reasoning68.3%
τ²-Bench - Agentic Tasks59.9%
TerminalBench - System Control36.4%
Specs
Context window
262Ktokens
Input pricing
$0.15per 1M tokens
Output pricing
$0.40per 1M tokens
Cached input
$0.14per 1M tokens

Prices in USD.

Key Capabilities & Ratings

Legal & E-Discovery90%
Agentic Actions90%
Financial Analysis90%
Tool Calling / API90%
Scientific Biology80%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderGoogleAnthropicAnthropic
Release DateApril 2, 2026July 24, 2026June 9, 2026
Knowledge Cutoff-May 2026-
Context & Limits
Context Window262K
1M
Best Context Window
1M
Best Context Window
Pricing (per 1M tokens)
Input Pricing
$0.15
Best Input Pricing
$5$10
Output Pricing
$0.40
Best Output Pricing
$25$50
Modalities
Inputs
imagetextvideo
textimage
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index22.3
63.1
Best Intelligence Index
62.1
Coding Index33.2
78.0
Best Coding Index
76.5
Agentic Index11.1
59.2
Best Agentic Index
56.6
Gemma 4 31B
Claude Opus 5
Claude Fable 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

76%
Gemma 4 31B
GPQA Benchmark
Score: 76%
Gemma 4 31B
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5
93%
Claude Fable 5
GPQA Benchmark
Score: 93%
Claude Fable 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

12%
Gemma 4 31B
Humanity's Last Exam
Score: 12%
Gemma 4 31B
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5
56%
Claude Fable 5
Humanity's Last Exam
Score: 56%
Claude Fable 5

Long Context Reasoning

Logical reasoning over long context windows.

43%
Gemma 4 31B
Long Context Reasoning
Score: 43%
Gemma 4 31B
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5
77%
Claude Fable 5
Long Context Reasoning
Score: 77%
Claude Fable 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Back to all models