AI Trend Notifier
EN
← wiki

$ cat wiki/models/muse-spark-1-2.md

Muse Spark 1.2

Compared with

Spec

AttributeValue
DeveloperMeta / Meta Superintelligence Labs (MSL)
Released2026-08-05
Announced2026-08-05
Context window1,000,000
Pricing$1.25/M input · $4.25/M output · $0.15/M cached input
Licenseunknown
AvailabilityMuse Code (beta, macOS + Linux), Meta Model API
License is unknown rather than proprietary: no source read states a licence for
1.2, and Muse Spark 1.1's terms are a sibling's, not this
release's. No weights were released
(source).

Release Date

2026-08-05, alongside Muse Code, the terminal coding agent it powers (source).

Benchmarks

Meta's own reported figures, against Muse Spark 1.1:

BenchmarkMuse Spark 1.1Muse Spark 1.2Gain
Terminal-Bench 2.176.2%82.9%+6.7
DeepSWE v1.153.0%59.3%+6.3
Meta publishes its methodology, and the caveats are Meta's own: the model was run
inside Meta's own vendor agent product (Muse Code), in **isolated Daytona cloud
sandboxes**, under an internal Meta evaluation framework, scored **pass@1 averaged
over five attempts** across the 89 tasks of the official Terminal-Bench 2.1 release.
Meta states that its agent tools and system prompts **"may not be specifically tuned for
proprietary third-party models"**, i.e. that competitors may not be measured at their
best (source).

In Meta's own comparison charts, Claude tops all three, with Muse Spark 1.2 second on the benchmarks Meta chose to highlight (source).

Independently measured by Artificial Analysis, which reports a 3-point gain on its Intelligence Index concentrated in agentic evaluations:

MetricMuse Spark 1.1Muse Spark 1.2
GDPval-AA v21371 Elo1631 Elo
Terminal-Bench v2.178%80%
τ³-Banking25%27%
This repo holds no local Artificial Analysis snapshot carrying these columns, so the
figures are quoted as reported and have nothing local to check them against
(source).

Use Cases

  • Muse Code — terminal coding agent, beta on macOS and Linux. Generates code and verifies it, and manages several persistent background sub-agents across a session (source).
  • Code generation, complex debugging, codebase understanding, end-to-end developer workflows — Meta's stated improvement areas over 1.1 (source).
  • Direct API use via the Meta Model API, unchanged in price and context from 1.1 (source).

The persistent background sub-agents are the product claim worth separating from the model claim: an agent that keeps sub-agents alive across a session is a harness design, and the benchmark figures above were produced with that harness in the loop. Meta's methodology says as much.

Compared To

ModelPrice (in/out per Mtok)ContextTerminal-Bench 2.1
Muse Spark 1.2$1.25/$4.251M82.9% (Meta) · 80% (AA)
Muse Spark 1.1 (Muse Spark (1.0 / 1.1))$1.25/$4.251M76.2% (Meta) · 78% (AA)
Meta held price and context window constant across the upgrade, which puts the whole
of the release in the model and the harness rather than the commercial terms
(source).

Conflicting Reports

The Terminal-Bench and DeepSWE figures are reported three different ways, and the page body follows Meta's own.

SourceTerminal-Bench 2.1DeepSWE 1.1SWE-Bench Pro
Meta (self-reported)82.9%59.3%not reported
Artificial Analysis (independent)80%not reportednot reported
kingy.ai, attributing to Meta's report80.053.361.5
The third row disagrees with the first on both shared benchmarks while claiming the
same origin, and carries a SWE-Bench Pro figure no other source read has. The second row
is a different measurement, not a contradiction — a different harness should be
expected to produce a different number, and Meta says so itself. Recorded rather than
reconciled (source).

Open Questions

  • Does 1.2 share the Muse Spark 1.1 backbone, or is it a separate training run? Its relationship to Watermelon is unstated.
  • Parameter count, architecture and training compute — none published.
  • Is Muse Code available outside the US? The Meta Model API public preview was US-developer-only at 1.1.
  • What licence governs 1.2, and does it differ from 1.1's?

Referenced by

Sources