# Anthropic Releases Claude Opus 5, Holding Price Flat While Claiming a Coding-Benchmark Jump

The company reports 43.3 percent on the Frontier-Bench v0.1 agentic coding test, more than double Opus 4.8, with a one-million-token context window and thinking enabled by default.

- Published: 2026-07-25T05:43:03.754Z
- Canonical: https://polylog.news/ai/2026-07-25/anthropic-releases-claude-opus-5-holding-price-flat-while-cl
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Anthropic](https://www.anthropic.com/news/claude-opus-5), [VentureBeat](https://venturebeat.com/orchestration/anthropic-launches-claude-opus-5-a-cheaper-ai-model-for-coding-agents-and-enterprise-workflows), [Polylog editors](https://polylog.news)

Anthropic on July 24 released [Claude Opus 5](https://www.anthropic.com/news/claude-opus-5), positioning it as a major advance for its top tier aimed at long-running agents, coding, and professional work. The pricing is unchanged from the prior Opus generation at $5 per million input tokens and $25 per million output tokens. The model ships with a one-million-token context window as both the default and the maximum. Extended reasoning is now on by default.

The main claims concern code and agents. Anthropic reports Opus 5 scoring [43.3 percent on Frontier-Bench v0.1](https://venturebeat.com/orchestration/anthropic-launches-claude-opus-5-a-cheaper-ai-model-for-coding-agents-and-enterprise-workflows), an agentic terminal-coding benchmark, against 18.7 percent for Opus 4.8, and 70.57 percent on OSWorld 2.0 for computer use. On the abstract-reasoning test ARC-AGI-3, third-party and vendor figures put Opus 5 at [30.2 percent](https://t.me/aipost/7619), well above the roughly 7.8 percent attributed to the next-best system and about 1.5 percent for Opus 4.8.

These are vendor-reported numbers on benchmarks that are new and, in the case of Frontier-Bench v0.1, controlled by parties with a stake in the result. None have been independently reproduced. The size of the ARC-AGI-3 gap warrants caution until outside labs run the test, because a twentyfold increase over a six-month-old model more often indicates a benchmark that rewards a specific agent setup than a genuine gain in reasoning.

Russian-language coverage from [AI ML Big Data](https://t.me/ai_machinelearning_big_data/10587) repeated Anthropic's description of Opus 5 as state-of-the-art for software engineering and agentic tasks. The flat pricing is the more concrete signal for engineers. Capability is being added at the same token cost, continuing the pattern of each Opus generation delivering more per dollar rather than charging a premium.

## What this means

Anthropic is competing on coding and agent reliability at a fixed price, which pressures OpenAI and Google to match its capability without raising per-token cost. The exposed parties are the closed-API rivals that sell on coding benchmarks and the enterprise buyers standardizing their agent systems on whichever model leads software-engineering-style leaderboards. The channel is distribution through developer tools. Claude Code and terminal agents entrench workflows, so a durable lead on agentic coding converts directly into higher API volume.

## What to watch

- Whether independent third parties can reproduce the Frontier-Bench and ARC-AGI-3 scores, which will show whether the gap is a genuine reasoning gain or a product of Anthropic's agent setup.
- Whether OpenAI and Google respond with a coding-focused release at matched or lower pricing, which would confirm price-flat capability gains as the new competitive baseline.
- Real-world agent failure rates on long-horizon tasks, where a one-million-token context helps only if the model stays coherent across it.
