Morning Edition · Friday, July 31, 2026Published at 1:48 AM EDT · New York
Anthropic's Claude Opus 5 Posts 96 Percent on SWE-bench Verified at Unchanged Opus Pricing
The new Opus scores 79.2 percent on the harder SWE-bench Pro variant, up from 69.2 percent for Opus 4.8, and ranks first on Anthropic's own agentic-coding evaluation.

Anthropic's Claude Opus 5, released July 24, is positioned as a significant improvement for the Opus tier on long-running agents, coding, and computer use, at the same price as the prior generation. Independent write-ups report it at 96.0 percent on SWE-bench Verified and 79.2 percent on SWE-bench Pro, against 69.2 percent for its predecessor Opus 4.8 on the Pro variant.
The gains are largest in the areas where competition among coding models is most intense. On SWE-bench Pro, Opus 5 ranks third overall, within a point of the leaders, while on Anthropic's own agentic-coding evaluation, Frontier-Bench, it ranks first. Buyers should read the second claim with the usual caution attached to a vendor-designed benchmark. A lab that publishes both the model and the evaluation controls the framing, and Frontier-Bench has no independent reproduction yet.
SWE-bench Verified scores are now approaching their maximum across frontier labs, which is why the harder Pro and senior-engineer variants are becoming the real discriminators. Opus 5's honest difference is incremental on the standard benchmark and larger on the harder one, delivered without a price increase.
What this means
Coding is the area where frontier labs now concentrate their competitive effort, and holding capability while keeping price unchanged is itself a competitive move against OpenAI's price cuts and cheaper open-weight coders. Anthropic gains through the developer-agent channel, where Claude Code and long-horizon task reliability drive lock-in more than a single benchmark number. The exposure for rivals is that a saturating SWE-bench Verified pushes the whole field toward harder, less reproducible evaluations that labs design themselves.
What to watch
- Independent results on SWE-bench Pro and senior-engineer leaderboards, which are harder to manipulate than the near-saturated Verified set.
- Whether third parties can reproduce Frontier-Bench rankings, the test of whether a lab-authored evaluation reflects real capability or favorable framing.
Observations to monitor, not financial advice.
Synthesized from: Anthropic · MarkTechPost
Part of a tracked trend
Frontier Labs Race on AI Coding Capability
Coding is becoming a primary competitive battleground among frontier labs, with incumbents standing up permanent coding teams and investing in new training stages (e.g. midtraining) to match leaders like Anthropic; expect recurring reorganizations, benchmarks, and model releases aimed specifically at code.
More from this edition
- OpenAI Cuts GPT-5.6 Luna Prices About 80 Percent, Turning the Model Race Into a Price War
- US Regulator Bars New Foreign-Made Humanoid Robots, Citing Supply-Chain Security
- Google DeepMind Ships Gemini Robotics ER 2 as a Reasoning Layer for Multi-Robot Tasks
- Google DeepMind Disbands the Original AlphaFold Team and Folds Science Into Gemini
- Meta Turns Muse Spark Into a Paid API and Zuckerberg Argues for Faster, Not Slower, AI
- GPTZero Flags Fabricated Citations in Four PwC Middle East Reports
- Preprint Probes Why Reinforcement Learning Beats Supervised Tuning on Math Reasoning
- Study Finds LLM Agents Deceive More Under Hidden, Conflicting Objectives
- New Jailbreak Uses Dual-Layer Encoding to Reconstruct Blocked Prompts Past Moderation
- ClinLens Benchmark Pushes Coding Agents Toward Long-Horizon Clinical Data Science
- ByteDance Readies Seedance 2.5 With Single-Pass 30-Second Video and Region-Level Editing