Granite 4.2 3B
IBM Granite 4.2 3B is a reasoning model for edge deployment, with thinking, non-thinking and low-effort modes, reasoning-augmented tool calling and 128K
Model Overview
Capabilities, design details, and architectural traits
Granite 4.2 3B - IBM's compact reasoning model for edge deployment
Granite 4.2 3B is the smallest member of IBM's Granite 4.2 family, the first Granite release built as dense reasoning models. It is positioned as a compact reasoning model optimized for edge deployment and resource-constrained environments, while sharing the same architecture and training pipeline as the larger 8B and 30B variants.
What sets it apart
| Trait | Detail |
|---|---|
| Edge-focused reasoning | Compact 3B model that still performs native chain-of-thought reasoning inside thinking tags before answering |
| Flexible thinking modes | Switch between thinking (default), non-thinking, and low-effort modes per query via chat-template parameters, balancing depth against latency |
| Reasoning-augmented tool calling | Reasons about which tools to invoke and why before making the call, producing more accurate function calls for agentic workflows |
| 128K context window | Natively supports a 128K token context window across the family |
| Open and trusted release | Apache 2.0 license with cryptographic signatures, ISO certification, and full transparency disclosures for unrestricted commercial and academic use |
| Multilingual dialog | Tested across 12 languages including English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, and Chinese |
How it was built
Granite 4.2 3B was pre-trained from scratch on roughly 15T tokens using a five-phase strategy, then supervised fine-tuned on chain-of-thought, reasoning, and agentic-trajectory data, followed by a multi-stage reinforcement learning pipeline. Unlike the 8B and 30B variants, the 3B model does not go through the additional agentic RL block, though it still supports native tool calling and emits tool calls in the OpenAI function-calling format when served through an OpenAI-compatible endpoint such as vLLM.
The result is a small model designed to bring explicit reasoning, flexible deliberation levels, and tool use to environments where compute is limited.
Benchmark Performance
Independent evaluations · Artificial Analysis
Accuracy & Capability Details
Compare Models Side-by-Side
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | IBM | Anthropic | Anthropic |
| Release Date | August 25, 2026 | September 22, 2026 | September 28, 2026 |
| Knowledge Cutoff | - | - | Jun 2026 |
| Context & Limits | |||
| Context Window | 524K | 1M Best Context Window | 1M Best Context Window |
| Pricing (per 1M tokens) | |||
| Input Pricing | $0.03 Best Input Pricing | $4 | $2 |
| Output Pricing | $0.12 Best Output Pricing | $20 | $10 |
| Modalities | |||
| Inputs | text | textimagefile | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 9.1 | 57.6 Best Intelligence Index | 56.0 |
| Coding Index | 17.5 | - | - |
| Agentic Index | 0.9 | - | - |
Humanity's Last Exam
Extremely difficult logical reasoning and knowledge.
Long Context Reasoning
Logical reasoning over long context windows.
SciCode Benchmark
Scientific coding and mathematical modeling.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.
Explore more from IBM
Other models by IBM
Top AI Models
Leading alternatives by intelligence score