Morning Edition · Saturday, July 25, 2026Published at 1:43 AM EDT · New York
The company reports 43.3 percent on the Frontier-Bench v0.1 agentic coding test, more than double Opus 4.8, with a one-million-token context window and thinking enabled by default.

Anthropic on July 24 released Claude Opus 5, positioning it as a major advance for its top tier aimed at long-running agents, coding, and professional work. The pricing is unchanged from the prior Opus generation at $5 per million input tokens and $25 per million output tokens. The model ships with a one-million-token context window as both the default and the maximum. Extended reasoning is now on by default.
The main claims concern code and agents. Anthropic reports Opus 5 scoring 43.3 percent on Frontier-Bench v0.1, an agentic terminal-coding benchmark, against 18.7 percent for Opus 4.8, and 70.57 percent on OSWorld 2.0 for computer use. On the abstract-reasoning test ARC-AGI-3, third-party and vendor figures put Opus 5 at 30.2 percent, well above the roughly 7.8 percent attributed to the next-best system and about 1.5 percent for Opus 4.8.
These are vendor-reported numbers on benchmarks that are new and, in the case of Frontier-Bench v0.1, controlled by parties with a stake in the result. None have been independently reproduced. The size of the ARC-AGI-3 gap warrants caution until outside labs run the test, because a twentyfold increase over a six-month-old model more often indicates a benchmark that rewards a specific agent setup than a genuine gain in reasoning.
Russian-language coverage from AI ML Big Data repeated Anthropic's description of Opus 5 as state-of-the-art for software engineering and agentic tasks. The flat pricing is the more concrete signal for engineers. Capability is being added at the same token cost, continuing the pattern of each Opus generation delivering more per dollar rather than charging a premium.
Anthropic, whose API revenue and valuation ride on a durable coding-benchmark lead that pressures OpenAI and Google to match capability without raising per-token price.
Part of a tracked trend
Frontier Labs Race on AI Coding Capability
Coding is becoming a primary competitive battleground among frontier labs, with incumbents standing up permanent coding teams and investing in new training stages (e.g. midtraining) to match leaders like Anthropic; expect recurring reorganizations, benchmarks, and model releases aimed specifically at code.
Start a discussion in Townsquare.
More from this edition
The release and flat pricing are independently confirmed, but the 43.3 percent Frontier-Bench and 30.2 percent ARC-AGI-3 figures are vendor-reported on new benchmarks that interested parties control and no outside lab has reproduced.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
Anthropic is competing on coding and agent reliability at a fixed price, which pressures OpenAI and Google to match its capability without raising per-token cost. The exposed parties are the closed-API rivals that sell on coding benchmarks and the enterprise buyers standardizing their agent systems on whichever model leads software-engineering-style leaderboards. The channel is distribution through developer tools. Claude Code and terminal agents entrench workflows, so a durable lead on agentic coding converts directly into higher API volume.
What to watch
Observations to monitor, not financial advice.
Synthesized from: Anthropic · VentureBeat · Polylog editors
Comments
0No comments yet.