# Anthropic Ships Claude Opus 5, Claiming 96 Percent on SWE-bench Verified

The company positions Opus 5 for long-running agents and coding, though on the harder SWE-bench Pro variant it trails two rival models by under a point.

- Published: 2026-07-29T05:45:43.018Z
- Canonical: https://polylog.news/ai/2026-07-29/anthropic-ships-claude-opus-5-claiming-96-percent-on-swe-ben
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Anthropic](https://www.anthropic.com/news/claude-opus-5)

Anthropic released [Claude Opus 5](https://www.anthropic.com/news/claude-opus-5) on July 24, describing it as a significant advance for its top Opus tier aimed at long-running agents, coding, and professional work. Third-party benchmark aggregators report the model at 96.0 percent on SWE-bench Verified, which would place it at the top of that leaderboard. It is priced at 5 dollars per million input tokens and 25 dollars per million output tokens with a 1-million-token context window.

  The results are more mixed on the harder [SWE-bench Pro](https://codingfleet.com/blog/swe-bench-pro-leaderboard-2026/) variant, where reported coverage puts Opus 5 near 79 percent, third behind two competing models by less than a point, while still well ahead of its predecessor Opus 4.8 at about 69 percent. On the newer Frontier-Bench and ARC-AGI-3 tests, aggregators report Opus 5 leading, though these are early and lightly reproduced.

  The numbers so far come from vendor materials and third-party aggregators rather than independent replication, and SWE-bench Verified is close to saturation at the top, which narrows the difference between frontier models to fractions of a point. The more lasting signal is price and context. Matching or exceeding prior frontier coding scores at the same token pricing continues the pattern of falling cost per unit of capability.

## What this means

Anthropic is defending the coding and agent segment where it leads, and the channel is capability plus switching cost through Claude Code and the API. Enterprises that standardize agent workflows on Opus gain a higher near-frontier capability limit at unchanged pricing, which pressures OpenAI, Google, and xAI to compete on both score and cost. With SWE-bench Verified near saturation, the competitive axis shifts to harder benchmarks like SWE-bench Pro and to real long-horizon agent reliability, where independent reproduction still lags the vendor claims.

## What to watch

- Independent SWE-bench Pro and long-horizon agent reproductions from third parties, which would confirm or undercut the 96 percent headline figure.
- Whether OpenAI and Google respond within weeks with priced frontier coding updates, a sign the coding race is now measured in fractions of a point.
