Polylog
The Polylog AI Intelligence Brief

Morning Edition · Friday, July 31, 2026Published at 1:48 AM EDT · New York

Anthropic's Claude Opus 5 Posts 96 Percent on SWE-bench Verified at Unchanged Opus Pricing

The new Opus scores 79.2 percent on the harder SWE-bench Pro variant, up from 69.2 percent for Opus 4.8, and ranks first on Anthropic's own agentic-coding evaluation.

Anthropic's Claude Opus 5 Posts 96 Percent on SWE-bench Verified at Unchanged Opus Pricing

Anthropic's Claude Opus 5, released July 24, is positioned as a significant improvement for the Opus tier on long-running agents, coding, and computer use, at the same price as the prior generation. Independent write-ups report it at 96.0 percent on SWE-bench Verified and 79.2 percent on SWE-bench Pro, against 69.2 percent for its predecessor Opus 4.8 on the Pro variant.

The gains are largest in the areas where competition among coding models is most intense. On SWE-bench Pro, Opus 5 ranks third overall, within a point of the leaders, while on Anthropic's own agentic-coding evaluation, Frontier-Bench, it ranks first. Buyers should read the second claim with the usual caution attached to a vendor-designed benchmark. A lab that publishes both the model and the evaluation controls the framing, and Frontier-Bench has no independent reproduction yet.

SWE-bench Verified scores are now approaching their maximum across frontier labs, which is why the harder Pro and senior-engineer variants are becoming the real discriminators. Opus 5's honest difference is incremental on the standard benchmark and larger on the harder one, delivered without a price increase.

What this means

Coding is the area where frontier labs now concentrate their competitive effort, and holding capability while keeping price unchanged is itself a competitive move against OpenAI's price cuts and cheaper open-weight coders. Anthropic gains through the developer-agent channel, where Claude Code and long-horizon task reliability drive lock-in more than a single benchmark number. The exposure for rivals is that a saturating SWE-bench Verified pushes the whole field toward harder, less reproducible evaluations that labs design themselves.

What to watch

  • Independent results on SWE-bench Pro and senior-engineer leaderboards, which are harder to manipulate than the near-saturated Verified set.
  • Whether third parties can reproduce Frontier-Bench rankings, the test of whether a lab-authored evaluation reflects real capability or favorable framing.

Observations to monitor, not financial advice.

2 sources

Synthesized from: Anthropic · MarkTechPost

Part of a tracked trend

Frontier Labs Race on AI Coding Capability

Coding is becoming a primary competitive battleground among frontier labs, with incumbents standing up permanent coding teams and investing in new training stages (e.g. midtraining) to match leaders like Anthropic; expect recurring reorganizations, benchmarks, and model releases aimed specifically at code.