Morning Edition · Monday, September 7, 2026Published at 2:22 AM EDT · New York
Fable 5.1 scores 52.6% on Terminal-Bench-Science 0.1, up from 24.7% for Fable 5, while the price of cached input drops from $1.00 to $0.25 per million tokens with headline pricing unchanged.

Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1, three months after Fable 5. The figure Anthropic emphasizes most is Terminal-Bench-Science 0.1, an agentic benchmark for scientific research tasks run in a terminal, where Fable 5.1 reaches 52.6% against 24.7% for its predecessor. On Terminal-Bench 4.0, Fable 5.1 scores 55.8% and Mythos 5.1 scores 60.9%, according to benchmark summaries of Anthropic's published figures.
The pricing change is what affects deployment costs immediately. List rates hold at $10 per million input tokens and $50 per million output tokens, but cache-read pricing falls 75%, from $1.00 to $0.25 per million tokens. VentureBeat reports that this translates to roughly 25% savings on typical workloads and up to 45% on agentic ones, the workload type where a long system prompt and accumulated tool output are re-read on every turn.
Mythos 5.1 is not a larger model or a different architecture. Anthropic describes it as the same model as Fable 5.1 with more permissive safeguards, available only through its trusted-access programs, and has not published separate pricing for it. The higher Terminal-Bench 4.0 score for Mythos is therefore a measure of what refusal behavior costs on agentic evaluations, not evidence of extra capability. Compared with OpenAI's GPT-6 Astra, which reports 57.9% on the same Terminal-Bench 4.0, the standard-safeguard Claude scores lower and the trusted-access variant scores higher, all on vendor-published numbers with no independent rerun yet.
Anthropic, which converts a cache-pricing cut into a lower cost per agent session while keeping its headline per-token rate unchanged for comparison purposes, and enterprise buyers running long agent loops.
Part of a tracked trend
Frontier Labs Race on AI Coding Capability
Coding is becoming a primary competitive battleground among frontier labs, with incumbents standing up permanent coding teams and investing in new training stages (e.g. midtraining) to match leaders like Anthropic; expect recurring reorganizations, benchmarks, and model releases aimed specifically at code.
Start a discussion in Townsquare.
More from this edition
Anthropic reports a standard error of roughly 3.5 to 4.5 points per model and scored some tasks at zero where production safeguards intervened, which the article's Fable-versus-Mythos comparison does not quantify, and no external party has rerun Terminal-Bench-Science 0.1.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
The cache-read cut targets the specific cost structure of agent loops, where the same context is billed repeatedly across turns, so it lowers the marginal cost of long-running agents without changing the headline rate vendors are compared on. Anthropic gains on cost per agent session while keeping its position in the standard per-token price comparison, which pressures OpenAI and Google to compete on cache pricing rather than headline price. The Fable-versus-Mythos gap also gives buyers a number for the throughput cost of safety filtering, something enterprise procurement teams will start asking about explicitly.
What to watch
Observations to monitor, not financial advice.
Source: Anthropic
Comments
0No comments yet.