Morning Edition · Monday, August 3, 2026Published at 1:38 AM EDT · New York
Anthropic's Opus 5 Matches Its Own Flagship on Coding at Half the Cost per Task
The company reports Opus 5 more than doubles Opus 4.8 on a software-engineering evaluation and reaches near-Fable 5 coding scores using a fraction of the reasoning tokens.

Anthropic's Claude Opus 5, released July 24, is the clearest recent evidence that capability per unit of inference compute is still improving. Anthropic says Opus 5 outperforms its prior models on Frontier-Bench, a software-engineering evaluation, and more than doubles the score of Opus 4.8 at a lower cost per task, while pricing stays at the Opus tier of roughly $5 per million input tokens and $25 per million output tokens.
The efficiency figures are what stand out. On the company's internal trading benchmark, Opus 5 reaches its result using about a seventh of the reasoning tokens and under half the latency of Opus 4.8. On CursorBench 3.2 at maximum effort, Anthropic reports that it comes within half a percentage point of Fable 5, its top model, at half the cost. Anthropic also claims Opus 5 scored roughly three times the next-best model on ARC-AGI 3, a test of novel problems, a figure that has not yet been independently reproduced.
These are vendor-reported results on a mix of public and internal evaluations, so the caution that applies to Alibaba applies here too. The difference is that Anthropic named the benchmarks and the token and latency differences, which are the numbers that matter to anyone paying for long-running agent runs, where token count rather than headline accuracy determines the cost.
Opus 5 is Anthropic's fourth model release in roughly two months, a pace that reflects how central coding and agentic work have become to the competition.
What this means
What is exposed here is inference-cost structure. When a model reaches near-flagship coding quality using a seventh of the reasoning tokens, the marginal cost of agentic workloads drops sharply. That pressures competitors selling similar quality at higher token counts and rewards whoever reduces tokens per task fastest. Anthropic gains on the measure that decides agent economics, which is cost per completed task rather than the highest benchmark score.
What to watch
- Independent reproduction of the ARC-AGI 3 and Frontier-Bench claims, since a threefold gap on novel-problem solving would be a real capability difference rather than a benchmark artifact.
- Whether rivals respond by cutting per-token prices or by publishing their own tokens-per-task figures, which would signal that inference efficiency has become the primary marketing focus.
Observations to monitor, not financial advice.
Synthesized from: Anthropic News · Anthropic News (hard questions)
Part of a tracked trend
Frontier Labs Race on AI Coding Capability
Coding is becoming a primary competitive battleground among frontier labs, with incumbents standing up permanent coding teams and investing in new training stages (e.g. midtraining) to match leaders like Anthropic; expect recurring reorganizations, benchmarks, and model releases aimed specifically at code.
More from this edition
- Alibaba Ships Qwen3.8-Max and Claims It Trails Only Anthropic's Top Model, Without Publishing the Numbers
- Berkshire's $339 Billion Treasury Position Is the Bear Case on AI Capex That Buffett Won't Say Directly
- A 6,000-Line C Engine Claims to Run the Full Kimi K3 Weights on a 64-Gigabyte Laptop
- A New Paper Names the Networking Bottleneck No Disaggregated Inference System Solves Correctly
- Researchers Propose a Pipeline That Uses Language Models to Generate and Validate Mathematical Conjectures
- A Benchmark Study Asks Whether AI Can Judge the Quality of AI-Generated Research
- Study Finds 40 Percent of Top TikTok Health Videos Are AI-Generated, Rising to 84 Percent for 'Health Tips' Searches
- Meta Opens a Paid Frontier API With Muse Spark 1.1, Ending Its Open-Only Posture
- Meta Puts Segment Anything and DINO Into National-Lab Science Projects
- Paper Proposes Cross-Model Auditing to Harden LLM Judges Against Their Own Biases
- Researchers Show Wallet Transaction-Simulation Previews Can Be Spoofed to Phish Crypto Users