Granite 4.2 8B
IBM Granite 4.2 8B is a mid-size dense reasoning model with chain-of-thought, flexible thinking modes, and reasoning-augmented tool calling under Apache 2.0.
Model Overview
Capabilities, design details, and architectural traits
Granite 4.2 8B - Mid-size dense reasoning model with flexible thinking modes
Granite 4.2 8B is the balanced, general-purpose member of IBM's Granite 4.2 family of dense reasoning language models. It performs step-by-step chain-of-thought reasoning before answering, and lets the user switch reasoning depth on a per-query basis within a single model.
| Trait | Detail |
|---|---|
| Built-in reasoning | Native chain-of-thought inside `` tags, improving math, coding, and multi-step logic performance |
| Flexible thinking modes | Three modes in one model: thinking (default), non-thinking, and low-effort, selected via chat-template parameters to balance depth vs. latency |
| Reasoning-augmented tool calling | Reasons about which tools to invoke and why before making the call, producing more accurate function calls for agentic workflows |
| Context window | Natively supports 128K, with long-context extension to 512K |
| Enterprise positioning | Balanced reasoning model for general-purpose enterprise applications, post-trained from Granite-4.1-8B-Base |
| Open licensing | Apache 2.0 with cryptographic signatures, ISO certification, and full transparency disclosures |
| Multilingual | Tested across 12 languages including English, German, Spanish, French, Japanese, Arabic, Korean, and Chinese |
Thinking modes as a first-class control
Rather than shipping separate reasoning and non-reasoning variants, Granite 4.2 8B exposes reasoning depth as a runtime switch. Full thinking produces complete chain-of-thought, non-thinking gives a direct answer with no reasoning overhead, and low-effort applies brief reasoning for simpler queries.
Trained for agentic work
The supervised fine-tuning stage combined publicly available permissive datasets, internally generated synthetic data targeting reasoning and tool calling, agentic traces collected across diverse tasks, and curated human-authored data.
Benchmark Performance
Independent evaluations · Artificial Analysis
Accuracy & Capability Details
Compare Models Side-by-Side
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | IBM | Anthropic | Anthropic |
| Release Date | August 25, 2026 | September 22, 2026 | September 28, 2026 |
| Knowledge Cutoff | - | - | Jun 2026 |
| Context & Limits | |||
| Context Window | 131K | 1M Best Context Window | 1M Best Context Window |
| Pricing (per 1M tokens) | |||
| Input Pricing | $0.06 Best Input Pricing | $4 | $2 |
| Output Pricing | $0.25 Best Output Pricing | $20 | $10 |
| Modalities | |||
| Inputs | text | textimagefile | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 11.1 | 57.6 Best Intelligence Index | 56.0 |
| Coding Index | 22.4 | - | - |
| Agentic Index | 1.3 | - | - |
Humanity's Last Exam
Extremely difficult logical reasoning and knowledge.
Long Context Reasoning
Logical reasoning over long context windows.
SciCode Benchmark
Scientific coding and mathematical modeling.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.
Explore more from IBM
Other models by IBM
Top AI Models
Leading alternatives by intelligence score