Morning Edition · Wednesday, July 29, 2026Published at 1:45 AM EDT · New York
Anthropic Ships Claude Opus 5, Claiming 96 Percent on SWE-bench Verified
The company positions Opus 5 for long-running agents and coding, though on the harder SWE-bench Pro variant it trails two rival models by under a point.

Anthropic released Claude Opus 5 on July 24, describing it as a significant advance for its top Opus tier aimed at long-running agents, coding, and professional work. Third-party benchmark aggregators report the model at 96.0 percent on SWE-bench Verified, which would place it at the top of that leaderboard. It is priced at 5 dollars per million input tokens and 25 dollars per million output tokens with a 1-million-token context window.
The results are more mixed on the harder SWE-bench Pro variant, where reported coverage puts Opus 5 near 79 percent, third behind two competing models by less than a point, while still well ahead of its predecessor Opus 4.8 at about 69 percent. On the newer Frontier-Bench and ARC-AGI-3 tests, aggregators report Opus 5 leading, though these are early and lightly reproduced.
The numbers so far come from vendor materials and third-party aggregators rather than independent replication, and SWE-bench Verified is close to saturation at the top, which narrows the difference between frontier models to fractions of a point. The more lasting signal is price and context. Matching or exceeding prior frontier coding scores at the same token pricing continues the pattern of falling cost per unit of capability.
What this means
Anthropic is defending the coding and agent segment where it leads, and the channel is capability plus switching cost through Claude Code and the API. Enterprises that standardize agent workflows on Opus gain a higher near-frontier capability limit at unchanged pricing, which pressures OpenAI, Google, and xAI to compete on both score and cost. With SWE-bench Verified near saturation, the competitive axis shifts to harder benchmarks like SWE-bench Pro and to real long-horizon agent reliability, where independent reproduction still lags the vendor claims.
What to watch
- Independent SWE-bench Pro and long-horizon agent reproductions from third parties, which would confirm or undercut the 96 percent headline figure.
- Whether OpenAI and Google respond within weeks with priced frontier coding updates, a sign the coding race is now measured in fractions of a point.
Observations to monitor, not financial advice.
Source: Anthropic
Part of a tracked trend
Frontier Labs Race on AI Coding Capability
Coding is becoming a primary competitive battleground among frontier labs, with incumbents standing up permanent coding teams and investing in new training stages (e.g. midtraining) to match leaders like Anthropic; expect recurring reorganizations, benchmarks, and model releases aimed specifically at code.
More from this edition
- China's CXMT Closes 466 Percent Above IPO Price, Becoming the Most Valuable Company Listed on the Mainland
- Nvidia Commits 5 Billion Dollars to Sutskever's Safe Superintelligence, a Lab With No Product
- Musk Sets August 7 for a 1.5-Trillion-Parameter Grok 4.6, With a 2.1-Trillion Grok 4.7 to Follow
- Terence Tao Tells the Congress of Mathematicians the Field Faces a Crisis in Its Foundations
- Companies That Cut Staff for AI Are Rehiring, With More Than Half of Leaders Calling the Layoffs a Mistake
- A New Paper Proposes Sparse, Block-Denoising Diffusion to Cut Language-Model Inference Cost
- Kernel Forge Puts an LLM Agent to Work Writing and Optimizing CUDA Kernels
- Two Papers Argue AI Safety Guardrails Do Not Compose Into Real Oversight
- Study Asks Whether Models Fake Alignment Even When Nothing Is at Stake
- Google Expands Gemini API Managed Agents With a 3.6 Flash Model and Lifecycle Hooks
- Meta Opens a Paid Model API With Muse Spark 1.1, Following Its Muse Image and Video Models