Morning Edition · Tuesday, September 8, 2026Published at 2:18 AM EDT · New York
The model reports 52.6% on Terminal-Bench-Science against 24.7% for its predecessor, while the sticker price stays at $10 and $50 per million tokens and cached input drops to $0.25.

Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on 1 September, three months after Fable 5. On Terminal-Bench-Science 0.1, an agentic scientific-research evaluation, Fable 5.1 reports 52.6%, compared with 24.7% for Fable 5, 29.0% for Claude Opus 5 and 22.4% for GPT-5.6 Sol. On Terminal-Bench 4.0, the coding evaluation, it reports 55.8%, versus 42.0% for Fable 5 and 52.3% for Opus 5. Other reported figures include 73.4% on CursorBench 3.2.0, 60.9% on Humanity's Last Exam without tools and 65.0% with tools, and 41.7% on the strict setting of OSWorld 2.0.
Every one of those figures comes from the vendor. The Terminal-Bench-Science jump is the most notable one, and a doubling on a young evaluation still at version 0.1 fits the pattern of targeted training on that specific task distribution rather than a general improvement in capability. Independent replication has not yet appeared.
The pricing change is easier to verify and more immediately useful. Input and output prices stay at $10 and $50 per million tokens, but cache reads fall from $1.00 to $0.25 per million, which Anthropic says lowers cost by roughly 25% on typical workloads and by up to 45% on agentic ones. Mythos 5.1 is described as the same model with lighter safeguards, available only to organizations vetted through Anthropic's Cyber Verification and Life Sciences Verification programs, and it reaches 60.9% on Terminal-Bench 4.0 under those looser restrictions.
What this means
Cutting the price of cache reads rather than the headline token price targets exactly the workload shape that agent frameworks produce: a long, fixed system prompt and tool schema that gets replayed at every step. Teams running multi-step agents on Anthropic's models see the largest savings, and competitors whose pricing does not distinguish cached from fresh input lose on total cost for the same task even when their list price is lower. The gap between Fable and Mythos also puts a number on the cost of the added safeguards: the same underlying weights score nearly five points higher on coding once those restrictions are loosened.
What to watch
Part of a tracked trend
Frontier Labs Race on AI Coding Capability
Coding is becoming a primary competitive battleground among frontier labs, with incumbents standing up permanent coding teams and investing in new training stages (e.g. midtraining) to match leaders like Anthropic; expect recurring reorganizations, benchmarks, and model releases aimed specifically at code.
Start a discussion in Townsquare.
More from this edition
Observations to monitor, not financial advice.
Synthesized from: Anthropic News · Anthropic News
Comments
0No comments yet.