Morning Edition · Wednesday, September 2, 2026Published at 2:20 AM EDT · New York
The restricted sibling model, Mythos 5.1, scores 60.9 percent on the same benchmark and stays behind vetting for cybersecurity and life-sciences professionals.

Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1. On Terminal-Bench 4.0, the agentic coding benchmark, Fable 5.1 scores 55.8 percent and Mythos 5.1 reaches 60.9 percent, against 42.0 percent for Fable 5 and 37.3 percent for OpenAI's GPT-5.6 Sol. On Terminal-Bench-Science 0.1, the newer model reports 52.6 percent compared with 24.7 percent for its predecessor, a result Russian-language technical channels picked up within hours.
The pricing change is the part engineers will feel first. Base rates are unchanged at $10 per million input tokens and $50 per million output tokens, but cache reads now cost 75 percent less. Anthropic puts the effect at roughly 25 percent lower cost on typical workloads and up to 45 percent on heavily agentic ones, which favors long-running agents that re-read a large context on every step rather than one-shot chat traffic.
The two model names describe one underlying model. Anthropic says the gap between Fable 5.1 and Mythos 5.1 reflects where earlier and less precise cyber safeguards intervened, and it keeps Mythos behind its Cyber Verification and Life Sciences Verification programs while Fable is generally available to anyone with a Claude account. That is the same access architecture OpenAI applied to Astra, developed separately by each company.
Every number here is vendor-reported. Terminal-Bench is a public benchmark, so third-party runs are possible, but the doubling on the science variant in particular has not yet been reproduced outside Anthropic.
Anthropic gains developer traffic by cutting the dominant cost line for agentic coding tools, and it establishes safety tiering as a way to sell one trained model into both open and regulated markets.
Part of a tracked trend
Frontier Labs Race on AI Coding Capability
Coding is becoming a primary competitive battleground among frontier labs, with incumbents standing up permanent coding teams and investing in new training stages (e.g. midtraining) to match leaders like Anthropic; expect recurring reorganizations, benchmarks, and model releases aimed specifically at code.
Start a discussion in Townsquare.
More from this edition
The 55.8 and 60.9 percent Terminal-Bench 4.0 figures and the cache read cut from $1.00 to $0.25 per million tokens come from Anthropic and its launch materials, the comparison against OpenAI's GPT-5.6 Sol uses a harness Anthropic configured, and the savings estimates of 25 to 45 percent describe workloads Anthropic selected rather than a measured customer average.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
The competitive move is not the benchmark score, it is the cache read price. Agentic coding tools spend most of their tokens re-reading the same repository context, so a 75 percent cut on cached input compresses the unit economics of every product built on Claude and pressures rivals to match on the cache line rather than the headline rate. Anthropic also demonstrated that a single trained model can be sold twice, once open and once gated, with the difference set by safeguard policy. That gives labs a way to serve regulated demand without publishing the capability broadly, and it makes safety tiering a pricing and distribution instrument.
What to watch
Observations to monitor, not financial advice.
Synthesized from: Anthropic News · Polylog editors
Comments
0No comments yet.