Polylog
The Polylog AI Intelligence Brief

Morning Edition · Monday, August 3, 2026Published at 1:38 AM EDT · New York

Anthropic's Opus 5 Matches Its Own Flagship on Coding at Half the Cost per Task

The company reports Opus 5 more than doubles Opus 4.8 on a software-engineering evaluation and reaches near-Fable 5 coding scores using a fraction of the reasoning tokens.

Anthropic's Opus 5 Matches Its Own Flagship on Coding at Half the Cost per Task

Anthropic's Claude Opus 5, released July 24, is the clearest recent evidence that capability per unit of inference compute is still improving. Anthropic says Opus 5 outperforms its prior models on Frontier-Bench, a software-engineering evaluation, and more than doubles the score of Opus 4.8 at a lower cost per task, while pricing stays at the Opus tier of roughly $5 per million input tokens and $25 per million output tokens.

The efficiency figures are what stand out. On the company's internal trading benchmark, Opus 5 reaches its result using about a seventh of the reasoning tokens and under half the latency of Opus 4.8. On CursorBench 3.2 at maximum effort, Anthropic reports that it comes within half a percentage point of Fable 5, its top model, at half the cost. Anthropic also claims Opus 5 scored roughly three times the next-best model on ARC-AGI 3, a test of novel problems, a figure that has not yet been independently reproduced.

These are vendor-reported results on a mix of public and internal evaluations, so the caution that applies to Alibaba applies here too. The difference is that Anthropic named the benchmarks and the token and latency differences, which are the numbers that matter to anyone paying for long-running agent runs, where token count rather than headline accuracy determines the cost.

Opus 5 is Anthropic's fourth model release in roughly two months, a pace that reflects how central coding and agentic work have become to the competition.

What this means

What is exposed here is inference-cost structure. When a model reaches near-flagship coding quality using a seventh of the reasoning tokens, the marginal cost of agentic workloads drops sharply. That pressures competitors selling similar quality at higher token counts and rewards whoever reduces tokens per task fastest. Anthropic gains on the measure that decides agent economics, which is cost per completed task rather than the highest benchmark score.

What to watch

  • Independent reproduction of the ARC-AGI 3 and Frontier-Bench claims, since a threefold gap on novel-problem solving would be a real capability difference rather than a benchmark artifact.
  • Whether rivals respond by cutting per-token prices or by publishing their own tokens-per-task figures, which would signal that inference efficiency has become the primary marketing focus.

Observations to monitor, not financial advice.

2 sources

Synthesized from: Anthropic News · Anthropic News (hard questions)

Part of a tracked trend

Frontier Labs Race on AI Coding Capability

Coding is becoming a primary competitive battleground among frontier labs, with incumbents standing up permanent coding teams and investing in new training stages (e.g. midtraining) to match leaders like Anthropic; expect recurring reorganizations, benchmarks, and model releases aimed specifically at code.