Morning Edition · Thursday, September 3, 2026Published at 2:26 AM EDT · New York
The model jumped from 55.0 for version 1.2 on the same long-horizon coding benchmark, a gain large enough that the size of it is itself the reason for caution.

Meta Superintelligence Labs released Muse Spark 1.3 and made it the model behind Muse Code, Meta's terminal-based coding agent. On DeepSWE v1.1, a long-horizon agentic coding benchmark, Meta reports 75.4 for Muse Spark 1.3 against 74.0 for Claude Opus 5 and 73.0 for GPT-5.6 Sol. The Telegram channel AI Post carried the same ranking.
The interesting number is not the one-and-a-half point lead. It is the 20-point jump from Muse Spark 1.2, which scored 55.0. A gain that large inside one minor version usually reflects either a genuine change in post-training and reinforcement-learning recipe or a change in how the agent harness runs the benchmark, and Meta has not published enough detail to separate the two. Independent verification of the DeepSWE result is still outstanding.
Muse Spark 1.3 is a closed model. Meta lists it as proprietary with no weights release, priced at $1.25 per million input tokens and $4.25 per million output tokens on the standard endpoint, with a roughly one-million-token context and text, image and video input. A much cheaper contributor endpoint is available to developers who let Meta train on their data, which is a data-acquisition instrument dressed as a pricing tier.
Meta's own framing sets the bar it will be judged against. Mark Zuckerberg called the release the largest single jump the team has made on coding and agentic work. That is a claim about sustained multi-step task completion, which is exactly the property that degrades fastest when a model leaves a benchmark harness and meets a real repository.
Meta, which needs developer adoption for a closed coding model and a data-for-discount endpoint, and which gains if buyers price Anthropic's agentic-coding premium as no longer defensible.
Part of a tracked trend
Frontier Labs Race on AI Coding Capability
Coding is becoming a primary competitive battleground among frontier labs, with incumbents standing up permanent coding teams and investing in new training stages (e.g. midtraining) to match leaders like Anthropic; expect recurring reorganizations, benchmarks, and model releases aimed specifically at code.
Start a discussion in Townsquare.
More from this edition
The 75.4 figure is Meta's own run on Meta's own harness and is not identical to the DeepSWE leaderboard configuration, so the 20-point jump from version 1.2 could reflect either post-training gains or harness changes, and no third party has reproduced it.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
Meta is competing for the coding market with a closed model and a data-for-discount pricing tier, putting it in direct competition with Anthropic's Claude Code and OpenAI's coding surface rather than with open-weight Llama derivatives. The exposure runs two ways: if DeepSWE reproduces externally, Anthropic loses its clearest remaining differentiator in agentic coding and the pricing premium attached to it. If it does not reproduce, Meta's benchmark credibility takes the damage, and the contributor endpoint starts to look like the actual product. Note who benefits from the claim being believed: Meta needs developer adoption more than it needs the revenue right now.
What to watch
Observations to monitor, not financial advice.
Synthesized from: Meta AI · Polylog editors · Meta Developer
Comments
0No comments yet.