Muse Spark 1.3
Meta's Muse Spark 1.3 flagship model for agentic coding: long single-thread workflows, clarifying questions, and confirmation before consequential actions.
Model Overview
Capabilities, design details, and architectural traits
Muse Spark 1.3 - Long-horizon agentic work in a single thread
Muse Spark 1.3 is Meta's flagship model for agentic workflows and coding, built to sustain open-ended objectives over one long conversation. Given a messy goal, it uses tools to generate its own context across conflicting sources, corrects gaps in its plan as it goes, and tracks what it has learned to deliver a final result. Meta frames the release as a step toward personal superintelligence.
| Trait | Detail |
|---|---|
| Single-thread multitasking | Juggles multiple workflows in one long thread and maps incoming prompts to the correct task, even when the user steers past requests or interrupts them mid-flow. |
| Active user collaboration | Asks clarifying questions when prompts are ambiguous, invokes help from the user when stuck, and confirms before taking consequential actions. |
| Self-calibration | Trained to know what it can and cannot do, and to surface hurdles instead of hallucinating outcomes. |
| Instruction fidelity | Follows complex, long-form instructions while preserving detailed requirements without dropping constraints or drifting from the requested workflow. |
| Adaptive reporting | Adapts to user preference on long tasks, either providing frequent updates or working silently in the background. |
| Harness-generalized training | Trained across a diverse set of harnesses so it generalizes to various agentic environments. |
| Efficiency | Used roughly 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2 in Meta engineer comparisons. |
Reasoning variants and inputs
The model ships with xhigh and max reasoning levels, with max arriving after additional safety testing. It keeps a 1M token context window and accepts text, image, and video input.
Where it runs
Muse Spark 1.3 is available in Muse Code, Meta's coding agent, and through the Meta Model API.
Benchmark Performance
Independent evaluations · Artificial Analysis
Accuracy & Capability Details
Compare Models Side-by-Side
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | Meta | Anthropic | Anthropic |
| Release Date | September 2, 2026 | September 1, 2026 | July 24, 2026 |
| Knowledge Cutoff | - | - | May 2026 |
| Context & Limits | |||
| Context Window | 1M | - | 1M |
| Pricing (per 1M tokens) | |||
| Input Pricing | Free Best Input Pricing | $10 | $5 |
| Output Pricing | Free Best Output Pricing | $50 | $25 |
| Modalities | |||
| Inputs | textimagefilevideo | textimagefile | textimage |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 62.1 | 65.7 Best Intelligence Index | 63.1 |
| Coding Index | 76.3 | 81.6 Best Coding Index | 78.0 |
| Agentic Index | 59.3 | 61.3 Best Agentic Index | 59.2 |
GPQA Benchmark
Graduate-level reasoning and expert Q&A evaluation.
Humanity's Last Exam
Extremely difficult logical reasoning and knowledge.
Long Context Reasoning
Logical reasoning over long context windows.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.