# Anthropic Ships Claude Opus 5 at Flat Pricing, Claiming Double Its Prior Agent Score

The model keeps the $5/$25 per-million-token price of Opus 4.8, but its own safety card reports it breached enterprise networks in 8 of 10 government red-team tests.

- Published: 2026-07-26T05:32:56.369Z
- Canonical: https://polylog.news/ai/2026-07-26/anthropic-ships-claude-opus-5-at-flat-pricing-claiming-doubl
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Anthropic News](https://www.anthropic.com/news/claude-opus-5), [Polylog editors](https://polylog.news)

Anthropic released [Claude Opus 5](https://www.anthropic.com/news/claude-opus-5) on July 24, positioning the model around long-running autonomous agents and coding rather than plain chat quality. It ships with a 1-million-token context window and up to 128K tokens of output, and Anthropic kept pricing unchanged at $5 per million input tokens and $25 per million output tokens, the same as the intervening Opus 4.8 release.

The company frames its main claims as relative gains rather than absolute scores. Anthropic says Opus 5 scores roughly three times the next-best model on ARC-AGI-3 and [more than doubles Opus 4.8](https://www.vellum.ai/blog/claude-opus-5-benchmarks-explained) on its internal Frontier-Bench agent evaluation at a lower cost per task. Because performance rose while token pricing did not, the release is primarily a capability-per-dollar improvement.

Two caveats matter for anyone planning to deploy it. First, the benchmark numbers come from Anthropic's own runs, and independent reviewers have noted that several comparisons were published as chart images rather than side-by-side tables, with code-review vendor [CodeRabbit measuring about 39% precision](https://www.coderabbit.ai/blog/opus-5-model-review) on actionable comments and roughly four times as many low-value suggestions as its baseline. Second, Anthropic's own safety documentation reports that Opus 5 [compromised enterprise networks in 8 of 10 government red-team exercises](https://www.techtimes.com/articles/321549/20260725/claude-opus-5-hacked-enterprise-networks-8-10-government-tests-safety-card-shows.htm), a figure that contradicts the model's positioning as a trustworthy autonomous agent.

## What this means

Keeping the price unchanged while improving capability is Anthropic's main competitive move. The company is competing at the top of the Opus tier on cost per completed task rather than on any single evaluation, which pressures OpenAI and Google to match performance without raising list prices. The offensive-cyber result on the safety card is the more consequential signal for buyers, because a model marketed for unsupervised, long-horizon agent work that also succeeds at network intrusion in most red-team runs is exactly the deployment profile that regulators and enterprise security teams are moving to constrain.

## What to watch

- Independent reproductions of the ARC-AGI-3 and Frontier-Bench claims on public harnesses, which would confirm or disprove the "double the prior model" framing.
- Whether enterprise and government buyers gate Opus 5 agent deployments behind additional controls after the red-team disclosure, a sign that safety-card findings are now shaping procurement.
