# Anthropic Ships Claude Opus 5, Its Fourth Model in Two Months, at Unchanged Pricing

The company reports a more-than-doubling on its Frontier-Bench coding evaluation and roughly 30 percent on ARC-AGI-3, well above prior scores, at $5 and $25 per million input and output tokens.

- Published: 2026-07-28T05:47:21.225Z
- Canonical: https://polylog.news/ai/2026-07-28/anthropic-ships-claude-opus-5-its-fourth-model-in-two-months
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Anthropic](https://www.anthropic.com/news/claude-opus-5)

Anthropic released [Claude Opus 5](https://www.anthropic.com/news/claude-opus-5) on July 24, describing it as a major improvement for the Opus tier aimed at long-running agents, coding, and professional work. The company kept pricing at $5 per million input tokens and $25 per million output tokens, matching the prior Opus generation.

On agentic and reasoning evaluations, Anthropic reports Opus 5 scoring about 30 percent on ARC-AGI-3, compared with 7.8 percent for GPT-5.6 Sol and 1.5 percent for Opus 4.8, and roughly 71 percent on OSWorld 2.0 for computer use. On Frontier-Bench, an agentic terminal-coding evaluation that measures whether a model can build working software from specifications, Anthropic says Opus 5 more than doubles its predecessor's score and passes every competitor tested, including the company's own Fable 5 tier, at a lower token price.

These figures are largely vendor-reported, and ARC-AGI-3 is a new evaluation without an established independent baseline, so the multiples deserve caution until third parties reproduce them. The release pace is itself informative. Four models in two months places coding and agentic autonomy at the center of Anthropic's plans, and holding the price steady suggests the competition is now on capability per dollar rather than on top benchmark scores alone.

## What this means

Coding and long-horizon agent autonomy are the main axis of competition among frontier labs, and Anthropic is pricing capability gains at flat token cost to defend its distribution against both open weights and rivals. The exposed parties are competing closed labs and coding-agent startups, because a flagship-class model at Opus pricing draws agentic workloads toward Anthropic through cost, not just quality. The unresolved question is independent reproduction of the ARC-AGI-3 and Frontier-Bench multiples.

## What to watch

- Independent reproductions of the ARC-AGI-3 and OSWorld 2.0 scores, which decide whether the reasoning jump is real or an artifact of a fresh benchmark.
- Whether real coding-agent deployments show the doubled Frontier-Bench result translating into fewer failed multi-file tasks in production.
