Toolbit.aiToolbit.ai

Find, compare, and explore the best AI tools to match your specific tasks and use cases.

Explore

  • AI Search
  • Compare ToolsNew
  • Browse Categories
  • Trending Tools
  • Most Popular
  • New Additions

Resources

  • Updates HubNew
  • AI News
  • ModelsNew
  • Blog Articles
  • NewsletterNew

Company

  • Launch a Tool
  • Advertise with Us
  • Guest Post
  • Contact Us
© 2026 Toolbit.ai. All rights reserved.
Privacy PolicyTerms & ConditionsDisclaimer
There's An AI For That favicon
There's An AI For That•The front page of AI for everyone
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Toolbit.ai
Toolbit.ai
Toolbit.ai
Toolbit.ai
UpdatesNew
Blog
Sign in
  1. Home
  2. Updates
  3. Models
  4. Mistral Large 3
Mistral
Released December 2, 2025

Mistral Large 3

Mistral Large 3 by Mistral AI is a frontier multimodal model with 256K context, vision input, native tool calling, and Apache 2.0 licensing.

Visit MistralAnnouncement
Inputs
Text
Image
File
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Mistral Large 3 - Open-Weight Multimodal Mixture-of-Experts Model

Mistral Large 3 is Mistral AI's open-weight, general-purpose multimodal model built on a granular Mixture-of-Experts architecture. It pairs a 673B-parameter MoE language model with a 2.5B-parameter vision encoder, totaling 41B active and 675B total parameters, and is released under the Apache 2.0 license.

Architecture

  • Granular MoE language model (673B total, 39B active) combined with a 2.5B-parameter vision encoder.
  • Trained from scratch on 3,000 NVIDIA H200 GPUs.

Context & Deployment

  • Supports a 256k-token context window.
  • Available in FP8 (single 8xB200/H200 node) and NVFP4 (single 8xH100/A100 node) quantized formats, with BF16 for multi-node deployment.
  • An EAGLE speculator checkpoint is provided for speculative decoding.

Multimodal & Tool Use

  • Accepts image and text input; documented guidance recommends keeping images near a 1:1 aspect ratio.
  • Offers native function calling and structured JSON outputs, intended for agentic use with a limited, well-defined toolset.

Benchmark Performance

Independent evaluations · Artificial Analysis

15.9%
Intelligence
20.1%
Coding Index
5.5%
Agentic Index

Accuracy & Capability Details

GPQA - Graduate Science68.0%
Humanity's Last Exam4.2%
SciCode - Scientific Coding36.2%
Instruction Following36.2%
Long Context Reasoning34.7%
τ²-Bench - Agentic Tasks24.6%
TerminalBench - System Control15.9%
Specs
Context window
262Ktokens
Input pricing
$0.50per 1M tokens
Output pricing
$1.50per 1M tokens

Prices in USD.

Key Capabilities & Ratings

Advanced Math80%
Translation & Lang80%
General Knowledge70%
Logical Logic70%
Creative Writing60%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderMistralAnthropicAnthropic
Release DateDecember 2, 2025July 24, 2026June 9, 2026
Knowledge Cutoff-May 2026-
Context & Limits
Context Window262K
1M
Best Context Window
1M
Best Context Window
Pricing (per 1M tokens)
Input Pricing
$0.50
Best Input Pricing
$5$10
Output Pricing
$1.50
Best Output Pricing
$25$50
Modalities
Inputs
textimagefile
textimage
textimagefile
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index15.9
63.1
Best Intelligence Index
62.1
Coding Index20.1
78.0
Best Coding Index
76.5
Agentic Index5.5
59.2
Best Agentic Index
56.6
Mistral Large 3
Claude Opus 5
Claude Fable 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

68%
Mistral Large 3
GPQA Benchmark
Score: 68%
Mistral Large 3
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5
93%
Claude Fable 5
GPQA Benchmark
Score: 93%
Claude Fable 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

4%
Mistral Large 3
Humanity's Last Exam
Score: 4%
Mistral Large 3
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5
56%
Claude Fable 5
Humanity's Last Exam
Score: 56%
Claude Fable 5

Long Context Reasoning

Logical reasoning over long context windows.

35%
Mistral Large 3
Long Context Reasoning
Score: 35%
Mistral Large 3
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5
77%
Claude Fable 5
Long Context Reasoning
Score: 77%
Claude Fable 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.

Back to all models