$ cat wiki/models/muse-spark-1-2.md
Muse Spark 1.2
Compared with
Spec
| Attribute | Value |
|---|---|
| Developer | Meta / Meta Superintelligence Labs (MSL) |
| Released | 2026-08-05 |
| Announced | 2026-08-05 |
| Context window | 1,000,000 |
| Pricing | $1.25/M input · $4.25/M output · $0.15/M cached input |
| License | unknown |
| Availability | Muse Code (beta, macOS + Linux), Meta Model API |
License is unknown rather than proprietary: no source read states a licence for | |
| 1.2, and Muse Spark 1.1's terms are a sibling's, not this | |
| release's. No weights were released | |
| (source). |
Release Date
2026-08-05, alongside Muse Code, the terminal coding agent it powers (source).
Benchmarks
Meta's own reported figures, against Muse Spark 1.1:
| Benchmark | Muse Spark 1.1 | Muse Spark 1.2 | Gain |
|---|---|---|---|
| Terminal-Bench 2.1 | 76.2% | 82.9% | +6.7 |
| DeepSWE v1.1 | 53.0% | 59.3% | +6.3 |
| Meta publishes its methodology, and the caveats are Meta's own: the model was run | |||
| inside Meta's own vendor agent product (Muse Code), in **isolated Daytona cloud | |||
| sandboxes**, under an internal Meta evaluation framework, scored **pass@1 averaged | |||
| over five attempts** across the 89 tasks of the official Terminal-Bench 2.1 release. | |||
| Meta states that its agent tools and system prompts **"may not be specifically tuned for | |||
| proprietary third-party models"**, i.e. that competitors may not be measured at their | |||
| best (source). |
In Meta's own comparison charts, Claude tops all three, with Muse Spark 1.2 second on the benchmarks Meta chose to highlight (source).
Independently measured by Artificial Analysis, which reports a 3-point gain on its Intelligence Index concentrated in agentic evaluations:
| Metric | Muse Spark 1.1 | Muse Spark 1.2 |
|---|---|---|
| GDPval-AA v2 | 1371 Elo | 1631 Elo |
| Terminal-Bench v2.1 | 78% | 80% |
| τ³-Banking | 25% | 27% |
| This repo holds no local Artificial Analysis snapshot carrying these columns, so the | ||
| figures are quoted as reported and have nothing local to check them against | ||
| (source). |
Use Cases
- Muse Code — terminal coding agent, beta on macOS and Linux. Generates code and verifies it, and manages several persistent background sub-agents across a session (source).
- Code generation, complex debugging, codebase understanding, end-to-end developer workflows — Meta's stated improvement areas over 1.1 (source).
- Direct API use via the Meta Model API, unchanged in price and context from 1.1 (source).
The persistent background sub-agents are the product claim worth separating from the model claim: an agent that keeps sub-agents alive across a session is a harness design, and the benchmark figures above were produced with that harness in the loop. Meta's methodology says as much.
Compared To
| Model | Price (in/out per Mtok) | Context | Terminal-Bench 2.1 |
|---|---|---|---|
| Muse Spark 1.2 | $1.25/$4.25 | 1M | 82.9% (Meta) · 80% (AA) |
| Muse Spark 1.1 (Muse Spark (1.0 / 1.1)) | $1.25/$4.25 | 1M | 76.2% (Meta) · 78% (AA) |
| Meta held price and context window constant across the upgrade, which puts the whole | |||
| of the release in the model and the harness rather than the commercial terms | |||
| (source). |
Conflicting Reports
The Terminal-Bench and DeepSWE figures are reported three different ways, and the page body follows Meta's own.
| Source | Terminal-Bench 2.1 | DeepSWE 1.1 | SWE-Bench Pro |
|---|---|---|---|
| Meta (self-reported) | 82.9% | 59.3% | not reported |
| Artificial Analysis (independent) | 80% | not reported | not reported |
| kingy.ai, attributing to Meta's report | 80.0 | 53.3 | 61.5 |
| The third row disagrees with the first on both shared benchmarks while claiming the | |||
| same origin, and carries a SWE-Bench Pro figure no other source read has. The second row | |||
| is a different measurement, not a contradiction — a different harness should be | |||
| expected to produce a different number, and Meta says so itself. Recorded rather than | |||
| reconciled (source). |
Open Questions
- Does 1.2 share the Muse Spark 1.1 backbone, or is it a separate training run? Its relationship to Watermelon is unstated.
- Parameter count, architecture and training compute — none published.
- Is Muse Code available outside the US? The Meta Model API public preview was US-developer-only at 1.1.
- What licence governs 1.2, and does it differ from 1.1's?