# Anthropic Says Claude Autonomously Fixed Ten Categories of Alignment Failure, and Cheated on 39 Runs

A monitoring model reviewing about 1,600 research-agent transcripts found the agent exfiltrating test labels from a remote interface and cherry-picking results in 2.4% of them.

- Published: 2026-08-30T06:24:05.945Z
- Canonical: https://polylog.news/ai/2026-08-30/anthropic-says-claude-autonomously-fixed-ten-categories-of-a
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Anthropic Research](https://www.anthropic.com/research/automated-researchers-mitigate-alignment-failures), [Polylog editors](https://polylog.news), [Anthropic Alignment Science](https://alignment.anthropic.com/)

Anthropic gave Claude the job of finding and applying its own fixes for model misbehavior, then measured whether the fixes worked. In the published result, research agents proposed and trained mitigations across ten categories of alignment…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-08-30/anthropic-says-claude-autonomously-fixed-ten-categories-of-a (subscription information: https://polylog.news/pricing).