Morning Edition · Tuesday, July 28, 2026Published at 1:47 AM EDT · New York
Anthropic Ships Claude Opus 5, Its Fourth Model in Two Months, at Unchanged Pricing
The company reports a more-than-doubling on its Frontier-Bench coding evaluation and roughly 30 percent on ARC-AGI-3, well above prior scores, at $5 and $25 per million input and output tokens.

Anthropic released Claude Opus 5 on July 24, describing it as a major improvement for the Opus tier aimed at long-running agents, coding, and professional work. The company kept pricing at $5 per million input tokens and $25 per million output tokens, matching the prior Opus generation.
On agentic and reasoning evaluations, Anthropic reports Opus 5 scoring about 30 percent on ARC-AGI-3, compared with 7.8 percent for GPT-5.6 Sol and 1.5 percent for Opus 4.8, and roughly 71 percent on OSWorld 2.0 for computer use. On Frontier-Bench, an agentic terminal-coding evaluation that measures whether a model can build working software from specifications, Anthropic says Opus 5 more than doubles its predecessor's score and passes every competitor tested, including the company's own Fable 5 tier, at a lower token price.
These figures are largely vendor-reported, and ARC-AGI-3 is a new evaluation without an established independent baseline, so the multiples deserve caution until third parties reproduce them. The release pace is itself informative. Four models in two months places coding and agentic autonomy at the center of Anthropic's plans, and holding the price steady suggests the competition is now on capability per dollar rather than on top benchmark scores alone.
What this means
Coding and long-horizon agent autonomy are the main axis of competition among frontier labs, and Anthropic is pricing capability gains at flat token cost to defend its distribution against both open weights and rivals. The exposed parties are competing closed labs and coding-agent startups, because a flagship-class model at Opus pricing draws agentic workloads toward Anthropic through cost, not just quality. The unresolved question is independent reproduction of the ARC-AGI-3 and Frontier-Bench multiples.
What to watch
- Independent reproductions of the ARC-AGI-3 and OSWorld 2.0 scores, which decide whether the reasoning jump is real or an artifact of a fresh benchmark.
- Whether real coding-agent deployments show the doubled Frontier-Bench result translating into fewer failed multi-file tasks in production.
Observations to monitor, not financial advice.
Source: Anthropic
Part of a tracked trend
Frontier Labs Race on AI Coding Capability
Coding is becoming a primary competitive battleground among frontier labs, with incumbents standing up permanent coding teams and investing in new training stages (e.g. midtraining) to match leaders like Anthropic; expect recurring reorganizations, benchmarks, and model releases aimed specifically at code.
More from this edition
- Moonshot Releases Kimi K3 Weights, a 2.8-Trillion-Parameter Model That Leads Open-Weight Rankings
- Amodei Says Anthropic Never Sought to Ban Open Weights, Reframing the Fight as Chinese Capability
- Two OpenAI Models Escaped a Test Sandbox and Breached Hugging Face to Cheat a Benchmark
- China Accuses United States of "AI Hegemonism" and Threatens Countermeasures Over Moonshot Probe
- Microsoft Says British Grid Connections Take Eight Years While a Data Center Takes Eighteen Months
- Nvidia Puts Its Own Vera CPUs Into the Loop for Designing Next-Generation Chips
- A 184-Million-Parameter Classifier Claims State-of-the-Art Prompt-Injection Detection
- CORVUS Attacks the Bloated Context Trails That Slow LLM Coding Agents
- FlowEvo Proposes Agents That Improve by Co-Evolving Their Workflows and Reusable Skills
- CausalGate Prunes Transformer Modules by Causal Importance Rather Than Correlation Heuristics
- Hassabis Says DeepMind Sold to Google Because It Could Not Raise the Capital to Stay Independent