Polylog
The Polylog AI Intelligence Brief

Morning Edition · Sunday, August 2, 2026Published at 1:44 AM EDT · New York

Anthropic's Claude Opus 5 Doubles Its Coding-Benchmark Score at Unchanged Pricing

Opus 5 scores 43.3% on Frontier-Bench versus Opus 4.8's 21.1% and comes within 0.5% of Claude Fable 5 on CursorBench at half the per-task cost.

Anthropic's Claude Opus 5 Doubles Its Coding-Benchmark Score at Unchanged Pricing

Anthropic shipped Claude Opus 5 on July 24, positioning it for long-running agents and professional coding work. On the company's published numbers, Opus 5 scores 43.3% on Frontier-Bench, more than double Opus 4.8's 21.1%, and at maximum effort on CursorBench 3.2 it comes within 0.5% of Claude Fable 5, the firm's top model, while costing roughly half as much per task.

Anthropic kept pricing flat at $5 and $25 per million input and output tokens, with a faster mode at about twice the price for roughly 2.5 times the speed. The model leads most of Anthropic's own benchmark suite, including GDPval-AA, OSWorld 2.0, and AutomationBench, and the company reports it scoring several times higher than the next model on ARC-AGI-3, a test of novel-problem solving.

These are vendor-reported figures on Anthropic-selected benchmarks, so the meaningful signal is the price-performance ratio rather than any single score. A model that reaches near-top-tier coding results at half the per-task cost of its top model shows how quickly the coding tier is being commoditized inside a single lab, which is what independent harness tests on real repositories will now test.

What this means

The channel is coding capability per dollar. By holding Opus pricing flat while closing most of the gap to its own top model, Anthropic compresses the price-performance frontier for agentic coding, which pressures rivals to match on cost, not just capability. The exposed parties are labs whose coding-tier models cost more for similar output, and developers building on coding agents gain a cheaper high-capability option, provided independent harnesses confirm the benchmark differences on real codebases.

What to watch

  • Independent SWE-bench-style results on real repositories, which test whether the doubled benchmark score translates into fewer failed agent runs.
  • Competing coding-tier releases and repricing from OpenAI, Google, and Chinese labs, which would show competition in the coding tier intensifying.

Observations to monitor, not financial advice.

2 sources

Synthesized from: Anthropic · MarkTechPost

Part of a tracked trend

Frontier Labs Race on AI Coding Capability

Coding is becoming a primary competitive battleground among frontier labs, with incumbents standing up permanent coding teams and investing in new training stages (e.g. midtraining) to match leaders like Anthropic; expect recurring reorganizations, benchmarks, and model releases aimed specifically at code.