# Anthropic Releases Claude Fable 5.1 at 55.8 Percent on Terminal-Bench 4.0 and Cuts Cache Read Prices by 75 Percent

The restricted sibling model, Mythos 5.1, scores 60.9 percent on the same benchmark and stays behind vetting for cybersecurity and life-sciences professionals.

- Published: 2026-09-02T06:20:39.906Z
- Canonical: https://polylog.news/ai/2026-09-02/anthropic-releases-claude-fable-5-1-at-55-8-percent-on-termi
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Anthropic News](https://www.anthropic.com/claude-fable-and-mythos-5-1), [Polylog editors](https://polylog.news)

Anthropic [released Claude Fable 5.1 and Claude Mythos 5.1](https://www.anthropic.com/claude-fable-and-mythos-5-1) on September 1. On Terminal-Bench 4.0, the agentic coding benchmark, Fable 5.1 scores 55.8 percent and Mythos 5.1 reaches 60.9 percent, against 42.0 percent for Fable 5 and 37.3 percent for OpenAI's GPT-5.6 Sol. On Terminal-Bench-Science 0.1, the newer model reports 52.6 percent compared with 24.7 percent for its predecessor, [a result Russian-language technical channels picked up within hours](https://t.me/ai_machinelearning_big_data/10811).

The pricing change is the part engineers will feel first. Base rates are unchanged at $10 per million input tokens and $50 per million output tokens, but cache reads now cost 75 percent less. Anthropic puts the effect at roughly 25 percent lower cost on typical workloads and [up to 45 percent on heavily agentic ones](https://the-decoder.com/anthropics-claude-fable-5-1-promises-better-coding-and-research-at-up-to-45-percent-less/), which favors long-running agents that re-read a large context on every step rather than one-shot chat traffic.

The two model names describe one underlying model. Anthropic says the gap between Fable 5.1 and Mythos 5.1 reflects where earlier and less precise cyber safeguards intervened, and it keeps Mythos behind its Cyber Verification and Life Sciences Verification programs while [Fable is generally available to anyone with a Claude account](https://t.me/aipost/8017). That is the same access architecture OpenAI applied to Astra, developed separately by each company.

Every number here is vendor-reported. Terminal-Bench is a public benchmark, so third-party runs are possible, but the doubling on the science variant in particular has not yet been reproduced outside Anthropic.

## What this means

The competitive move is not the benchmark score, it is the cache read price. Agentic coding tools spend most of their tokens re-reading the same repository context, so a 75 percent cut on cached input compresses the unit economics of every product built on Claude and pressures rivals to match on the cache line rather than the headline rate. Anthropic also demonstrated that a single trained model can be sold twice, once open and once gated, with the difference set by safeguard policy. That gives labs a way to serve regulated demand without publishing the capability broadly, and it makes safety tiering a pricing and distribution instrument.

## What to watch

- Whether independent Terminal-Bench 4.0 runs by outside groups land near 55.8 percent. A gap between vendor and third-party numbers would say more about harness configuration than about model capability.
- Whether OpenAI or Google respond by cutting cached-input pricing rather than by shipping a new model. That would confirm the cost of context, not raw quality, is the live battleground for coding agents.
- How many organizations clear Anthropic's Cyber Verification Program. A wide approval list would mean gating is light, and a narrow one would mean capability access is concentrating in a small set of firms.
