# Anthropic's Claude Opus 5 Posts 96 Percent on SWE-bench Verified at Unchanged Opus Pricing

The new Opus scores 79.2 percent on the harder SWE-bench Pro variant, up from 69.2 percent for Opus 4.8, and ranks first on Anthropic's own agentic-coding evaluation.

- Published: 2026-07-31T05:48:01.029Z
- Canonical: https://polylog.news/ai/2026-07-31/anthropic-s-claude-opus-5-posts-96-percent-on-swe-bench-veri
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Anthropic](https://www.anthropic.com/news/claude-opus-5), [MarkTechPost](https://www.marktechpost.com/2026/07/24/meet-the-new-claude-opus-5-frontier-class-agentic-coding-and-computer-use-at-unchanged-opus-pricing/)

Anthropic's [Claude Opus 5](https://www.anthropic.com/news/claude-opus-5), released July 24, is positioned as a significant improvement for the Opus tier on long-running agents, coding, and computer use, at the same price as the prior generation. Independent write-ups report it at [96.0 percent on SWE-bench Verified and 79.2 percent on SWE-bench Pro](https://www.marktechpost.com/2026/07/24/meet-the-new-claude-opus-5-frontier-class-agentic-coding-and-computer-use-at-unchanged-opus-pricing/), against 69.2 percent for its predecessor Opus 4.8 on the Pro variant.

The gains are largest in the areas where competition among coding models is most intense. On SWE-bench Pro, Opus 5 ranks third overall, within a point of the leaders, while on Anthropic's own agentic-coding evaluation, Frontier-Bench, it ranks first. Buyers should read the second claim with the usual caution attached to a vendor-designed benchmark. A lab that publishes both the model and the evaluation controls the framing, and Frontier-Bench has no independent reproduction yet.

SWE-bench Verified scores are now approaching their maximum across frontier labs, which is why the harder Pro and senior-engineer variants are becoming the real discriminators. Opus 5's honest difference is incremental on the standard benchmark and larger on the harder one, delivered without a price increase.

## What this means

Coding is the area where frontier labs now concentrate their competitive effort, and holding capability while keeping price unchanged is itself a competitive move against OpenAI's price cuts and cheaper open-weight coders. Anthropic gains through the developer-agent channel, where Claude Code and long-horizon task reliability drive lock-in more than a single benchmark number. The exposure for rivals is that a saturating SWE-bench Verified pushes the whole field toward harder, less reproducible evaluations that labs design themselves.

## What to watch

- Independent results on SWE-bench Pro and senior-engineer leaderboards, which are harder to manipulate than the near-saturated Verified set.
- Whether third parties can reproduce Frontier-Bench rankings, the test of whether a lab-authored evaluation reflects real capability or favorable framing.
