Morning Edition · Sunday, September 6, 2026Published at 2:32 AM EDT · New York
The company reports 52.6% on Terminal-Bench-Science against 24.7% for Fable 5, and ships a second variant with lighter safeguards restricted to vetted United States organizations.

Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1, three months after Fable 5. The headline result is on Terminal-Bench-Science 0.1, a benchmark that scores a model running scientific research tasks in a terminal. Anthropic reports Fable 5.1 at 52.6%, against 24.7% for Fable 5, 29.0% for Claude Opus 5 and 22.4% for OpenAI's GPT-5.6 Sol.
The company also published the standard error, between roughly 3.5 and 4.5 points per model. That is an unusual disclosure, and it matters for the conclusion: a 28-point gap over Fable 5 survives that uncertainty comfortably, while several of the narrower comparisons Anthropic draws elsewhere do not. On Terminal-Bench 4.0 the numbers are 55.8% for Fable 5.1 and 60.9% for Mythos 5.1, and on Humanity's Last Exam, Fable 5.1 reaches 60.9% without tools and 65.0% with them, per MarkTechPost's summary of the release.
Mythos 5.1 is the more consequential product decision. Anthropic describes it as the same underlying model with lighter safeguards, available only to organizations that clear its Cyber Verification Program or Life Sciences Verification Program, and only in the United States at launch. The five-point Terminal-Bench gap between the two variants is a direct measurement of how much capability the safety layer costs.
Anthropic also cut cache read pricing by 75%, which matters more for agentic workloads than the benchmark deltas do. Long-running agents re-read the same context repeatedly, and cache economics decide whether a multi-hour research run is affordable.
Anthropic, which turns a safety-vetting program into a paid product tier, and the United States organizations cleared for Mythos 5.1 that get a measurably stronger model than any competitor outside the program can buy.
Part of a tracked trend
AI Moves Into Autonomous Scientific Discovery and Clinical Care
Over the next 3-9 months, AI systems move beyond text tasks into running real scientific experiments and managing clinical care, backed by peer-reviewed and benchmarked evidence of chemist- and physician-level performance.
Start a discussion in Townsquare.
More from this edition
The scores match Anthropic's release and MarkTechPost's summary, and independent aggregation currently places Fable 5.1 ahead of GPT-6 Astra, but Terminal-Bench-Science 0.1 is a new benchmark run on the vendor's own harness with a stated standard error of 3.5 to 4.5 points, and the five-point gap between Fable and Mythos is a single measurement, not a general estimate of what the safety layer costs.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
Anthropic is converting safety vetting into a product tier, the same structure OpenAI adopted with its Daybreak program, and the measured five-point capability gap between the gated and ungated variants gives enterprises a concrete figure to weigh in negotiations. Contract research organizations, biotech firms and security teams inside the United States gain access to a materially stronger model than everyone else, which pushes non-United States labs toward open-weight alternatives where no vetting layer exists. The cache read price cut is the quieter competitive move, because it lowers the marginal cost of the long agent runs that Anthropic wants to dominate.
What to watch
Observations to monitor, not financial advice.
Source: Anthropic News
Comments
0No comments yet.