Morning Edition · Saturday, July 25, 2026Published at 1:43 AM EDT · New York
Anthropic Releases Claude Opus 5, Holding Price Flat While Claiming a Coding-Benchmark Jump
The company reports 43.3 percent on the Frontier-Bench v0.1 agentic coding test, more than double Opus 4.8, with a one-million-token context window and thinking enabled by default.

Anthropic on July 24 released Claude Opus 5, positioning it as a major advance for its top tier aimed at long-running agents, coding, and professional work. The pricing is unchanged from the prior Opus generation at $5 per million input tokens and $25 per million output tokens. The model ships with a one-million-token context window as both the default and the maximum. Extended reasoning is now on by default.
The main claims concern code and agents. Anthropic reports Opus 5 scoring 43.3 percent on Frontier-Bench v0.1, an agentic terminal-coding benchmark, against 18.7 percent for Opus 4.8, and 70.57 percent on OSWorld 2.0 for computer use. On the abstract-reasoning test ARC-AGI-3, third-party and vendor figures put Opus 5 at 30.2 percent, well above the roughly 7.8 percent attributed to the next-best system and about 1.5 percent for Opus 4.8.
These are vendor-reported numbers on benchmarks that are new and, in the case of Frontier-Bench v0.1, controlled by parties with a stake in the result. None have been independently reproduced. The size of the ARC-AGI-3 gap warrants caution until outside labs run the test, because a twentyfold increase over a six-month-old model more often indicates a benchmark that rewards a specific agent setup than a genuine gain in reasoning.
Russian-language coverage from AI ML Big Data repeated Anthropic's description of Opus 5 as state-of-the-art for software engineering and agentic tasks. The flat pricing is the more concrete signal for engineers. Capability is being added at the same token cost, continuing the pattern of each Opus generation delivering more per dollar rather than charging a premium.
- If true, who benefits
Anthropic, whose API revenue and valuation ride on a durable coding-benchmark lead that pressures OpenAI and Google to match capability without raising per-token price.
- The nuance
The release and flat pricing are independently confirmed, but the 43.3 percent Frontier-Bench and 30.2 percent ARC-AGI-3 figures are vendor-reported on new benchmarks that interested parties control and no outside lab has reproduced.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
Anthropic is competing on coding and agent reliability at a fixed price, which pressures OpenAI and Google to match its capability without raising per-token cost. The exposed parties are the closed-API rivals that sell on coding benchmarks and the enterprise buyers standardizing their agent systems on whichever model leads software-engineering-style leaderboards. The channel is distribution through developer tools. Claude Code and terminal agents entrench workflows, so a durable lead on agentic coding converts directly into higher API volume.
What to watch
- Whether independent third parties can reproduce the Frontier-Bench and ARC-AGI-3 scores, which will show whether the gap is a genuine reasoning gain or a product of Anthropic's agent setup.
- Whether OpenAI and Google respond with a coding-focused release at matched or lower pricing, which would confirm price-flat capability gains as the new competitive baseline.
- Real-world agent failure rates on long-horizon tasks, where a one-million-token context helps only if the model stays coherent across it.
Observations to monitor, not financial advice.
Synthesized from: Anthropic · VentureBeat · Polylog editors
Part of a tracked trend
Frontier Labs Race on AI Coding Capability
Coding is becoming a primary competitive battleground among frontier labs, with incumbents standing up permanent coding teams and investing in new training stages (e.g. midtraining) to match leaders like Anthropic; expect recurring reorganizations, benchmarks, and model releases aimed specifically at code.
More from this edition
- Researchers Say a Kimi K3 Agent Swarm Found Redis Code-Execution Flaws in 27 Minutes
- OpenAI and Apollo Research Publish a Method to Detect Hidden Reward-Seeking in Models
- South Korea Commits to Roughly 260,000 Nvidia GPUs for Sovereign AI
- Anthropic Doubles Its AI-Policy Donation to $40 Million Ahead of US Midterms
- Meta's Brain2Qwerty Decodes Typed Sentences From Non-Invasive Brain Scans at 61 Percent Word Accuracy
- Meta's Open Models Cut a Month of DOE Beamline Analysis to Minutes
- A Utah Copper Mine Adds Boston Dynamics Robots to a Fully Autonomous Operation
- MoE Interpretability Papers Probe How Expert Routing Encodes Knowledge and Frequency
- Meta Launches Muse Media Models Aimed at Editable, Production-Ready Output
- New Study Extracts LLMs' Implicit Theories of What Makes Writing Good
- Musk Says AI Will Soon Outstrip Humans by More Than the Human-Chimpanzee Gap