# Meta's Muse Spark 1.3 Scores 75.4 on DeepSWE, Above Claude Opus 5 and GPT-5.6 Sol

The model jumped from 55.0 for version 1.2 on the same long-horizon coding benchmark, a gain large enough that the size of it is itself the reason for caution.

- Published: 2026-09-03T06:26:17.363Z
- Canonical: https://polylog.news/ai/2026-09-03/meta-s-muse-spark-1-3-scores-75-4-on-deepswe-above-claude-op
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Meta AI](https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/), [Polylog editors](https://polylog.news), [Meta Developer](https://developer.meta.com/ai/models/muse-spark/)

Meta Superintelligence Labs released Muse Spark 1.3 and made it the model behind Muse Code, Meta's terminal-based coding agent. On DeepSWE v1.1, a long-horizon agentic coding benchmark, [Meta reports 75.4 for Muse Spark 1.3 against 74.0 for Claude Opus 5 and 73.0 for GPT-5.6 Sol](https://officechai.com/ai/muse-spark-1-3-benchmarks/). The Telegram channel AI Post [carried the same ranking](https://t.me/aipost/8031).

The interesting number is not the one-and-a-half point lead. It is the 20-point jump from Muse Spark 1.2, which scored 55.0. A gain that large inside one minor version usually reflects either a genuine change in post-training and reinforcement-learning recipe or a change in how the agent harness runs the benchmark, and Meta has not published enough detail to separate the two. Independent verification of the DeepSWE result is still outstanding.

Muse Spark 1.3 is a closed model. Meta lists it as proprietary with no weights release, priced at [$1.25 per million input tokens and $4.25 per million output tokens](https://www.explainx.ai/blog/meta-muse-spark-1-3-launch-benchmarks-pricing-september-2026) on the standard endpoint, with a roughly one-million-token context and text, image and video input. A much cheaper contributor endpoint is available to developers who let Meta train on their data, which is a data-acquisition instrument dressed as a pricing tier.

Meta's own framing sets the bar it will be judged against. Mark Zuckerberg called the release the largest single jump the team has made on coding and agentic work. That is a claim about sustained multi-step task completion, which is exactly the property that degrades fastest when a model leaves a benchmark harness and meets a real repository.

## What this means

Meta is competing for the coding market with a closed model and a data-for-discount pricing tier, putting it in direct competition with Anthropic's Claude Code and OpenAI's coding surface rather than with open-weight Llama derivatives. The exposure runs two ways: if DeepSWE reproduces externally, Anthropic loses its clearest remaining differentiator in agentic coding and the pricing premium attached to it. If it does not reproduce, Meta's benchmark credibility takes the damage, and the contributor endpoint starts to look like the actual product. Note who benefits from the claim being believed: Meta needs developer adoption more than it needs the revenue right now.

## What to watch

- Whether a third party runs DeepSWE v1.1 against Muse Spark 1.3 with a published harness configuration, since the 20-point jump from version 1.2 is the specific thing that needs replication.
- How many developers take the cheaper contributor endpoint, which would show that labs can now buy real coding traces with discounts instead of paying for annotation.
- Whether Meta releases weights for any Muse Spark model, because keeping the frontier tier closed while the Llama line stays open marks a permanent split in Meta's strategy.
