MBZUAI Institute of Foundation Models
Released September 3, 2026

K2 Horizon 375B A23B

K2 Horizon 375B A23B by the MBZUAI Institute of Foundation Models is a sparse MoE flagship with 23B active parameters for complex reasoning, coding and

Inputs
Text
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

K2 Horizon 375B-A23B - the enterprise powerhouse of the Horizon fleet

K2 Horizon 375B-A23B is a sparse mixture-of-experts flagship built for demanding enterprise workloads, and it anchors a fleet that spans from 0.9B edge models up to this deployment target.

TraitDetail
ArchitectureSparse MoE with 375 billion total parameters and roughly 23 billion active per token, drawing on the capacity of a much larger model without using every parameter for every token
PositioningDescribed by IFM as the fleet's enterprise powerhouse
Target workloadsComplex reasoning, software engineering, research and long-horizon agentic tasks
OpennessReleased under Apache 2.0 as part of what IFM calls the first fully open model fleet for agents, exposing the complete development process through agentic post-training
Released artifactsIntermediate checkpoints, training data or detailed data-construction recipes, training code, configurations, fine-grained logs, evaluation results and final weights
Deployment profileA multi-accelerator, cluster-scale deployment rather than a single-device model

One connected fleet

The 375B-A23B model shares core architecture, vocabulary, training methodology, interfaces, evaluation infrastructure and deployment tooling with its five fleet siblings. This consistency is designed to make it easier to move between sizes, route work dynamically and study capability across scale, so teams can prototype on a smaller model and scale up to the flagship.

Reasoning built into pretraining

Across the fleet, IFM incorporated reasoning directly into pretraining, with nearly 17% of the pretraining corpus consisting of explicit reasoning trajectories. The open release is intended to let researchers study how reasoning, tool use, planning and agentic capabilities emerge, and to reproduce or adapt the methods rather than treat the final checkpoint as an opaque starting point.

Benchmark Performance

Independent evaluations · Artificial Analysis

47.3%
Intelligence
61.5%
Coding Index
42.8%
Agentic Index

Accuracy & Capability Details

GPQA - Graduate Science87.3%
Humanity's Last Exam32.0%
SciCode - Scientific Coding40.7%
Long Context Reasoning75.7%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderMBZUAI Institute of Foundation ModelsAnthropicAnthropic
Release DateSeptember 3, 2026September 1, 2026July 24, 2026
Knowledge Cutoff--May 2026
Context & Limits
Context Window524K-
1M
Best Context Window
Pricing (per 1M tokens)
Input Pricing
Free
Best Input Pricing
$10$5
Output Pricing
Free
Best Output Pricing
$50$25
Modalities
Inputs
text
textimagefile
textimage
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index47.3
65.7
Best Intelligence Index
63.1
Coding Index61.5
81.6
Best Coding Index
78.0
Agentic Index42.8
61.3
Best Agentic Index
59.2
K2 Horizon 375B A23B
Claude Fable 5.1
Claude Opus 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

87%
K2 Horizon 375B A23B
GPQA Benchmark
Score: 87%
K2 Horizon 375B A23B
94%
Claude Fable 5.1
GPQA Benchmark
Score: 94%
Claude Fable 5.1
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

32%
K2 Horizon 375B A23B
Humanity's Last Exam
Score: 32%
K2 Horizon 375B A23B
59%
Claude Fable 5.1
Humanity's Last Exam
Score: 59%
Claude Fable 5.1
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5

Long Context Reasoning

Logical reasoning over long context windows.

76%
K2 Horizon 375B A23B
Long Context Reasoning
Score: 76%
K2 Horizon 375B A23B
80%
Claude Fable 5.1
Long Context Reasoning
Score: 80%
Claude Fable 5.1
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.