Polylog
The Polylog AI Intelligence Brief

Morning Edition · Sunday, July 26, 2026Published at 1:32 AM EDT · New York

Anthropic Ships Claude Opus 5 at Flat Pricing, Claiming Double Its Prior Agent Score

The model keeps the $5/$25 per-million-token price of Opus 4.8, but its own safety card reports it breached enterprise networks in 8 of 10 government red-team tests.

Anthropic Ships Claude Opus 5 at Flat Pricing, Claiming Double Its Prior Agent Score

Anthropic released Claude Opus 5 on July 24, positioning the model around long-running autonomous agents and coding rather than plain chat quality. It ships with a 1-million-token context window and up to 128K tokens of output, and Anthropic kept pricing unchanged at $5 per million input tokens and $25 per million output tokens, the same as the intervening Opus 4.8 release.

The company frames its main claims as relative gains rather than absolute scores. Anthropic says Opus 5 scores roughly three times the next-best model on ARC-AGI-3 and more than doubles Opus 4.8 on its internal Frontier-Bench agent evaluation at a lower cost per task. Because performance rose while token pricing did not, the release is primarily a capability-per-dollar improvement.

Two caveats matter for anyone planning to deploy it. First, the benchmark numbers come from Anthropic's own runs, and independent reviewers have noted that several comparisons were published as chart images rather than side-by-side tables, with code-review vendor CodeRabbit measuring about 39% precision on actionable comments and roughly four times as many low-value suggestions as its baseline. Second, Anthropic's own safety documentation reports that Opus 5 compromised enterprise networks in 8 of 10 government red-team exercises, a figure that contradicts the model's positioning as a trustworthy autonomous agent.

Veracity: Corroborated
80/100
If true, who benefits

Anthropic, which converts a red-team disclosure into evidence of safety rigor while advancing its capability-per-dollar positioning against OpenAI and Google, and the broader case for pre-emptive AI regulation that favors incumbent labs.

The nuance

The 8-of-10 figure is real but comes from a UK government red-team against a simulated, deliberately non-hardened network, and the same system card reports the model's lowest-ever misalignment score alongside elevated evaluation awareness, while the "double the prior model" gains rest on vendor-run, not independently reproduced, benchmarks.

An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.

What this means

Keeping the price unchanged while improving capability is Anthropic's main competitive move. The company is competing at the top of the Opus tier on cost per completed task rather than on any single evaluation, which pressures OpenAI and Google to match performance without raising list prices. The offensive-cyber result on the safety card is the more consequential signal for buyers, because a model marketed for unsupervised, long-horizon agent work that also succeeds at network intrusion in most red-team runs is exactly the deployment profile that regulators and enterprise security teams are moving to constrain.

What to watch

  • Independent reproductions of the ARC-AGI-3 and Frontier-Bench claims on public harnesses, which would confirm or disprove the "double the prior model" framing.
  • Whether enterprise and government buyers gate Opus 5 agent deployments behind additional controls after the red-team disclosure, a sign that safety-card findings are now shaping procurement.

Observations to monitor, not financial advice.

2 sources

Synthesized from: Anthropic News · Polylog editors

Part of a tracked trend

Frontier Model Efficiency Gains

Capability per unit of training and inference compute keeps improving, letting newer models match prior frontier performance far more cheaply and gradually loosening the link between raw scale and capability.