# Anthropic's Claude Fable 5.1 More Than Doubles Its Predecessor on an Agentic Science Benchmark

The company reports 52.6% on Terminal-Bench-Science against 24.7% for Fable 5, and ships a second variant with lighter safeguards restricted to vetted United States organizations.

- Published: 2026-09-06T06:32:46.692Z
- Canonical: https://polylog.news/ai/2026-09-06/anthropic-s-claude-fable-5-1-more-than-doubles-its-predecess
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Anthropic News](https://www.anthropic.com/claude-fable-and-mythos-5-1)

Anthropic released [Claude Fable 5.1 and Claude Mythos 5.1](https://www.anthropic.com/claude-fable-and-mythos-5-1) on September 1, three months after Fable 5. The headline result is on Terminal-Bench-Science 0.1, a benchmark that scores a model running scientific research tasks in a terminal. Anthropic reports Fable 5.1 at 52.6%, against 24.7% for Fable 5, 29.0% for Claude Opus 5 and 22.4% for OpenAI's GPT-5.6 Sol.

The company also published the standard error, between roughly 3.5 and 4.5 points per model. That is an unusual disclosure, and it matters for the conclusion: a 28-point gap over Fable 5 survives that uncertainty comfortably, while several of the narrower comparisons Anthropic draws elsewhere do not. On Terminal-Bench 4.0 the numbers are 55.8% for Fable 5.1 and 60.9% for Mythos 5.1, and on Humanity's Last Exam, Fable 5.1 reaches 60.9% without tools and 65.0% with them, [per MarkTechPost's summary of the release](https://www.marktechpost.com/2026/09/01/anthropic-releases-claude-fable-5-1-and-claude-mythos-5-1-52-6-on-terminal-bench-science-and-75-cheaper-cache-reads/).

Mythos 5.1 is the more consequential product decision. Anthropic describes it as the same underlying model with lighter safeguards, available only to organizations that clear its Cyber Verification Program or Life Sciences Verification Program, and only in the United States at launch. The five-point Terminal-Bench gap between the two variants is a direct measurement of how much capability the safety layer costs.

Anthropic also cut cache read pricing by 75%, which matters more for agentic workloads than the benchmark deltas do. Long-running agents re-read the same context repeatedly, and cache economics decide whether a multi-hour research run is affordable.

## What this means

Anthropic is converting safety vetting into a product tier, the same structure OpenAI adopted with its Daybreak program, and the measured five-point capability gap between the gated and ungated variants gives enterprises a concrete figure to weigh in negotiations. Contract research organizations, biotech firms and security teams inside the United States gain access to a materially stronger model than everyone else, which pushes non-United States labs toward open-weight alternatives where no vetting layer exists. The cache read price cut is the quieter competitive move, because it lowers the marginal cost of the long agent runs that Anthropic wants to dominate.

## What to watch

- Whether independent groups reproduce the Terminal-Bench-Science result, since the benchmark is new, and large jumps on new evaluations often shrink once tested with a different harness.
- How many organizations actually clear the Cyber Verification and Life Sciences Verification programs, because a small approved list makes the tier a signaling exercise rather than a real distribution channel.
- Whether OpenAI or Google match the cache read price cut, which would confirm that inference caching is now a competitive axis rather than a billing detail.
