# Anthropic's Claude Fable 5.1 More Than Doubles Its Score on an Agentic Science Benchmark and Cuts Cache Reads 75%

Fable 5.1 scores 52.6% on Terminal-Bench-Science 0.1, up from 24.7% for Fable 5, while the price of cached input drops from $1.00 to $0.25 per million tokens with headline pricing unchanged.

- Published: 2026-09-07T06:22:00.097Z
- Canonical: https://polylog.news/ai/2026-09-07/anthropic-s-claude-fable-5-1-more-than-doubles-its-score-on
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Anthropic](https://www.anthropic.com/claude-fable-and-mythos-5-1)

Anthropic released [Claude Fable 5.1 and Claude Mythos 5.1](https://www.anthropic.com/claude-fable-and-mythos-5-1) on September 1, three months after Fable 5. The figure Anthropic emphasizes most is Terminal-Bench-Science 0.1, an agentic benchmark for scientific research tasks run in a terminal, where Fable 5.1 reaches 52.6% against 24.7% for its predecessor. On Terminal-Bench 4.0, Fable 5.1 scores 55.8% and Mythos 5.1 scores 60.9%, according to [benchmark summaries of Anthropic's published figures](https://www.marktechpost.com/2026/09/01/anthropic-releases-claude-fable-5-1-and-claude-mythos-5-1-52-6-on-terminal-bench-science-and-75-cheaper-cache-reads/).

The pricing change is what affects deployment costs immediately. List rates hold at $10 per million input tokens and $50 per million output tokens, but cache-read pricing falls 75%, from $1.00 to $0.25 per million tokens. [VentureBeat reports](https://venturebeat.com/technology/anthropics-claude-fable-5-1-and-mythos-5-1-arrive-with-a-75-cost-reduction-for-fable-cache-reads) that this translates to roughly 25% savings on typical workloads and up to 45% on agentic ones, the workload type where a long system prompt and accumulated tool output are re-read on every turn.

Mythos 5.1 is not a larger model or a different architecture. Anthropic describes it as the same model as Fable 5.1 with more permissive safeguards, available only through its trusted-access programs, and has not published separate pricing for it. The higher Terminal-Bench 4.0 score for Mythos is therefore a measure of what refusal behavior costs on agentic evaluations, not evidence of extra capability. Compared with OpenAI's GPT-6 Astra, which [reports 57.9%](https://openai.com/index/gpt-6-astra/) on the same Terminal-Bench 4.0, the standard-safeguard Claude scores lower and the trusted-access variant scores higher, all on vendor-published numbers with no independent rerun yet.

## What this means

The cache-read cut targets the specific cost structure of agent loops, where the same context is billed repeatedly across turns, so it lowers the marginal cost of long-running agents without changing the headline rate vendors are compared on. Anthropic gains on cost per agent session while keeping its position in the standard per-token price comparison, which pressures OpenAI and Google to compete on cache pricing rather than headline price. The Fable-versus-Mythos gap also gives buyers a number for the throughput cost of safety filtering, something enterprise procurement teams will start asking about explicitly.

## What to watch

- Whether independent evaluators reproduce the 52.6% Terminal-Bench-Science figure, since the benchmark is new and the jump from 24.7% is large enough to invite scrutiny of the test setup.
- Whether OpenAI or Google respond by cutting cached-input pricing rather than base rates, which would confirm that cache pricing, not base pricing, is now the main way vendors compete on agent workloads.
- How widely Anthropic opens trusted access to Mythos, because a permissive-safeguard tier available to many customers represents a different safety posture than one available to only a few.
