Toolbit.aiToolbit.ai

Find, compare, and explore the best AI tools to match your specific tasks and use cases.

Explore

  • AI Search
  • Compare ToolsNew
  • Browse Categories
  • Trending Tools
  • Most Popular
  • New Additions

Resources

  • Updates HubNew
  • AI News
  • ModelsNew
  • Blog Articles
  • NewsletterNew

Company

  • Launch a Tool
  • Advertise with Us
  • Guest Post
  • Contact Us
© 2026 Toolbit.ai. All rights reserved.
Privacy PolicyTerms & ConditionsDisclaimer
There's An AI For That favicon
There's An AI For That•The front page of AI for everyone
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Toolbit.ai
Toolbit.ai
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Sign in
  1. Home
  2. Updates
  3. Models
  4. Gemma 3 4B
Google
Released March 12, 2025Cutoff August 2024

Gemma 3 4B

Gemma 3 4B by Google DeepMind is an open-weights multimodal model with 128K context, SigLIP pan-and-scan vision, 5:1 local/global attention, and 140+ language support.

Visit GoogleAnnouncement
Inputs
Text
Image
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Gemma 3 4B - Smallest Open-Weights Gemma 3 Model

Gemma 3 4B is a multimodal, open-weights model from Google DeepMind, built from the same research used to create Gemini. It is the smallest size in the Gemma 3 family to support image input - the 1B variant is text-only. Despite its size, it shares the full 128K-token context window with the 12B and 27B variants, not the 32K limit of the 1B. It targets deployment on resource-constrained hardware such as laptops and single GPUs.

Architecture and Vision Handling That Define the 4B

TraitDetail
5:1 local/global attention interleavingAlternates 5 local sliding window layers (1024-token window) per 1 global attention layer, optimizing KV-cache memory for long contexts
SigLIP vision encoder with pan-and-scanImages normalized to 896x896; pan-and-scan adaptively crops non-square or high-resolution images into tiles, each re-encoded separately to preserve detail
128K context windowShared with 12B and 27B; not available in the 1B size
Open weights in two variantsReleased as both a pre-trained base and an instruction-tuned (-it) version
Distillation-based post-trainingInstruction-tuned variant trained using knowledge distillation and reinforcement learning
Multilingual scopeSupports over 140 languages

Attention Design Specific to Gemma 3

The 5:1 local-to-global attention ratio is a documented change from Gemma 2, which used a 1:1 ratio with a 4096-token local window. In Gemma 3, the local window is reduced to 1024 tokens and the ratio shifts to 5:1, reducing KV-cache memory use without degrading language model perplexity. This pattern applies across the 4B, 12B, and 27B sizes.

Benchmark Performance

Independent evaluations · Artificial Analysis

1.0%
Intelligence
2.7%
Coding Index

Accuracy & Capability Details

GPQA - Graduate Science29.1%
Humanity's Last Exam5.3%
SciCode - Scientific Coding7.3%
Instruction Following28.3%
Long Context Reasoning6.7%
τ²-Bench - Agentic Tasks5.0%
TerminalBench - System Control0.8%
Specs
Context window
131Ktokens

Prices in USD.

Key Capabilities & Ratings

Legal & E-Discovery70%
Financial Analysis70%
Advanced Math60%
Agentic Actions60%
Scientific Biology60%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderGoogleAnthropicAnthropic
Release DateMarch 12, 2025July 24, 2026June 9, 2026
Knowledge CutoffAug 2024May 2026-
Context & Limits
Context Window131K
1M
Best Context Window
1M
Best Context Window
Pricing (per 1M tokens)
Input Pricing
Free
Best Input Pricing
$5$10
Output Pricing
Free
Best Output Pricing
$25$50
Modalities
Inputs
textimage
textimage
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index1.0
63.1
Best Intelligence Index
62.1
Coding Index2.7
78.0
Best Coding Index
76.5
Agentic Index-
59.2
Best Agentic Index
56.6
Gemma 3 4B
Claude Opus 5
Claude Fable 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

29%
Gemma 3 4B
GPQA Benchmark
Score: 29%
Gemma 3 4B
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5
93%
Claude Fable 5
GPQA Benchmark
Score: 93%
Claude Fable 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

5%
Gemma 3 4B
Humanity's Last Exam
Score: 5%
Gemma 3 4B
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5
56%
Claude Fable 5
Humanity's Last Exam
Score: 56%
Claude Fable 5

Long Context Reasoning

Logical reasoning over long context windows.

7%
Gemma 3 4B
Long Context Reasoning
Score: 7%
Gemma 3 4B
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5
77%
Claude Fable 5
Long Context Reasoning
Score: 77%
Claude Fable 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Back to all models