Meta
Released September 2, 2026

Muse Spark 1.3

Meta's Muse Spark 1.3 flagship model for agentic coding: long single-thread workflows, clarifying questions, and confirmation before consequential actions.

Inputs
Text
Image
File
Video
Outputs
Text

Model Overview

Capabilities, design details, and architectural traits

Muse Spark 1.3 - Long-horizon agentic work in a single thread

Muse Spark 1.3 is Meta's flagship model for agentic workflows and coding, built to sustain open-ended objectives over one long conversation. Given a messy goal, it uses tools to generate its own context across conflicting sources, corrects gaps in its plan as it goes, and tracks what it has learned to deliver a final result. Meta frames the release as a step toward personal superintelligence.

TraitDetail
Single-thread multitaskingJuggles multiple workflows in one long thread and maps incoming prompts to the correct task, even when the user steers past requests or interrupts them mid-flow.
Active user collaborationAsks clarifying questions when prompts are ambiguous, invokes help from the user when stuck, and confirms before taking consequential actions.
Self-calibrationTrained to know what it can and cannot do, and to surface hurdles instead of hallucinating outcomes.
Instruction fidelityFollows complex, long-form instructions while preserving detailed requirements without dropping constraints or drifting from the requested workflow.
Adaptive reportingAdapts to user preference on long tasks, either providing frequent updates or working silently in the background.
Harness-generalized trainingTrained across a diverse set of harnesses so it generalizes to various agentic environments.
EfficiencyUsed roughly 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2 in Meta engineer comparisons.

Reasoning variants and inputs

The model ships with xhigh and max reasoning levels, with max arriving after additional safety testing. It keeps a 1M token context window and accepts text, image, and video input.

Where it runs

Muse Spark 1.3 is available in Muse Code, Meta's coding agent, and through the Meta Model API.

Benchmark Performance

Independent evaluations · Artificial Analysis

62.1%
Intelligence
76.3%
Coding Index
59.3%
Agentic Index

Accuracy & Capability Details

GPQA - Graduate Science93.8%
Humanity's Last Exam49.1%
SciCode - Scientific Coding57.3%
Long Context Reasoning79.0%

Compare Models Side-by-Side

Evaluate specifications, pricing, and independent benchmark indices

Model Details
General Info
ProviderMetaAnthropicAnthropic
Release DateSeptember 2, 2026September 1, 2026July 24, 2026
Knowledge Cutoff--May 2026
Context & Limits
Context Window1M-1M
Pricing (per 1M tokens)
Input Pricing
Free
Best Input Pricing
$10$5
Output Pricing
Free
Best Output Pricing
$50$25
Modalities
Inputs
textimagefilevideo
textimagefile
textimage
Outputs
text
text
text
Benchmarks (0-100)
Intelligence Index62.1
65.7
Best Intelligence Index
63.1
Coding Index76.3
81.6
Best Coding Index
78.0
Agentic Index59.3
61.3
Best Agentic Index
59.2
Muse Spark 1.3
Claude Fable 5.1
Claude Opus 5

GPQA Benchmark

Graduate-level reasoning and expert Q&A evaluation.

94%
Muse Spark 1.3
GPQA Benchmark
Score: 94%
Muse Spark 1.3
94%
Claude Fable 5.1
GPQA Benchmark
Score: 94%
Claude Fable 5.1
93%
Claude Opus 5
GPQA Benchmark
Score: 93%
Claude Opus 5

Humanity's Last Exam

Extremely difficult logical reasoning and knowledge.

49%
Muse Spark 1.3
Humanity's Last Exam
Score: 49%
Muse Spark 1.3
59%
Claude Fable 5.1
Humanity's Last Exam
Score: 59%
Claude Fable 5.1
55%
Claude Opus 5
Humanity's Last Exam
Score: 55%
Claude Opus 5

Long Context Reasoning

Logical reasoning over long context windows.

79%
Muse Spark 1.3
Long Context Reasoning
Score: 79%
Muse Spark 1.3
80%
Claude Fable 5.1
Long Context Reasoning
Score: 80%
Claude Fable 5.1
76%
Claude Opus 5
Long Context Reasoning
Score: 76%
Claude Opus 5

Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.