IBM
Released April 29, 2026

Granite 4.1 8B

Granite 4.1 8B by IBM is a dense, decoder-only open-source model with 512K context, four-stage RL alignment, and tool calling for enterprise automation.

Inputs
Text
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Granite 4.1 8B - Dense Decoder-Only Enterprise Language Model

Granite 4.1 8B is IBM's balanced-tier open-source language model, part of the Granite 4.1 family. It uses a dense, decoder-only transformer architecture - deliberately avoiding sparse MoE routing - to deliver predictable latency and stable token usage for enterprise workloads. IBM documents the 8B instruct variant as consistently matching or outperforming the previous-generation Granite 4.0-H-Small.

Architecture

Granite-4.1-8B is built on a dense decoder-only transformer with: GQA (Grouped Query Attention), RoPE positional embeddings, MLP with SwiGLU activation, RMSNorm, and shared input/output embeddings. No sparse routing, no MoE layers.

Benchmark Performance

Independent evaluations · Artificial Analysis

6.6%
Intelligence
9.5%
Coding Index

Accuracy & Capability Details

GPQA - Graduate Science43.3%
Humanity's Last Exam3.8%
SciCode - Scientific Coding21.8%
Instruction Following38.6%
Long Context Reasoning13.3%
τ²-Bench - Agentic Tasks27.8%
TerminalBench - System Control0.0%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderIBMAnthropicAnthropic
Release DateApril 29, 2026September 22, 2026September 28, 2026
Knowledge Cutoff--Jun 2026
Context & Limits
Context Window524K
1M
Best Context Window
1M
Best Context Window
Pricing (per 1M tokens)
Input Pricing
$0.05
Best Input Pricing
$4$2
Output Pricing
$0.10
Best Output Pricing
$20$10
Modalities
Inputs
text
textimagefile
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index6.6
57.6
Best Intelligence Index
56.0
Coding Index9.5--
Agentic Index---
Granite 4.1 8B
Claude Opus 5.5
Claude Sonnet 5.5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

4%
Granite 4.1 8B
Humanity's Last Exam
Score: 4%
Granite 4.1 8B
61%
Claude Opus 5.5
Humanity's Last Exam
Score: 61%
Claude Opus 5.5
55%
Claude Sonnet 5.5
Humanity's Last Exam
Score: 55%
Claude Sonnet 5.5

Long Context Reasoning

Logical reasoning over long context windows.

13%
Granite 4.1 8B
Long Context Reasoning
Score: 13%
Granite 4.1 8B
85%
Claude Opus 5.5
Long Context Reasoning
Score: 85%
Claude Opus 5.5
83%
Claude Sonnet 5.5
Long Context Reasoning
Score: 83%
Claude Sonnet 5.5

SciCode Benchmark

Scientific coding and mathematical modeling.

22%
Granite 4.1 8B
SciCode Benchmark
Score: 22%
Granite 4.1 8B
67%
Claude Opus 5.5
SciCode Benchmark
Score: 67%
Claude Opus 5.5
61%
Claude Sonnet 5.5
SciCode Benchmark
Score: 61%
Claude Sonnet 5.5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Explore more from IBM

Other models by IBM

Top AI Models

Leading alternatives by intelligence score

View all