Muse Spark 1.2: coding model co-trained with Muse Code, with long-horizon training, self-improvement, and goal-conditioned planning for multi-step work.
Capabilities, design details, and architectural traits
Muse Spark 1.2 is Meta's coding-focused language model, designed specifically for long-running, multi-step software engineering across large repositories rather than short autocomplete suggestions. It is the model powering Muse Code, Meta's terminal coding agent.
| Trait | Detail |
|---|---|
| Co-training with Muse Code | Trained alongside the Muse Code agent using rejection-sampled harness trajectories and recipe optimizations for goals, compaction, and subagents, plus integration of the Muse Code toolset for harness compatibility |
| Self-improvement loop | Muse Spark 1.1 generated challenging coding environments and instruction-following templates, then graded candidate solutions to produce training data for 1.2 |
| Long-horizon training | Extensively trained on whole-repository generation, large end-to-end projects, and auto-research tasks |
| Goal-conditioned planning | Uses planning to sequence work, goal conditioning to maintain direction, and context compaction to retain knowledge across extended sessions |
The model's training pipeline is unusual in two ways. First, Muse Spark 1.2 was co-trained with Muse Code so the model performs best when paired with that specific agent runtime. Second, its predecessor Muse Spark 1.1 was used to generate and grade training environments, creating a self-improvement loop that improved instruction-following precision.
Independent evaluations · Artificial Analysis
Evaluate specifications, pricing, and independent benchmark indices
| Model Details | |||
|---|---|---|---|
| General Info | |||
| Provider | Meta | Anthropic | Anthropic |
| Release Date | August 5, 2026 | July 24, 2026 | June 9, 2026 |
| Knowledge Cutoff | - | May 2026 | - |
| Context & Limits | |||
| Context Window | - | 1M | 1M |
| Pricing (per 1M tokens) | |||
| Input Pricing | $1.25 Best Input Pricing | $5 | $10 |
| Output Pricing | $4.25 Best Output Pricing | $25 | $50 |
| Modalities | |||
| Inputs | text | textimage | textimagefile |
| Outputs | text | text | text |
| Benchmarks (0-100) | |||
| Intelligence Index | 56.8 | 63.1 Best Intelligence Index | 62.1 |
| Coding Index | 72.2 | 78.0 Best Coding Index | 76.5 |
| Agentic Index | 49.3 | 59.2 Best Agentic Index | 56.6 |
Graduate-level reasoning and expert Q&A evaluation.
Extremely difficult logical reasoning and knowledge.
Logical reasoning over long context windows.
Independent evaluation data provided by Artificial Analysis. To view the latest benchmarks and full details, visit their official site.