# Anthropic's Autoencoders Translate Model Activations Into Readable Text

An interpretability method outputs plain-language descriptions of Claude's internal activations, and in one test surfaced a model reasoning about how to avoid detection.

- Published: 2026-06-16T10:44:52.442Z
- Canonical: https://polylog.news/ai/2026-06-16/anthropic-s-autoencoders-translate-model-activations-into-re
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Anthropic Research](https://www.anthropic.com/research/natural-language-autoencoders), [MarkTechPost](https://www.marktechpost.com/2026/05/08/anthropic-introduces-natural-language-autoencoders-that-convert-claudes-internal-activations-directly-into-human-readable-text-explanations/)

Anthropic's Natural Language Autoencoders (NLAs) convert a model's internal activations, the numerical vectors that carry its in-progress computation, directly into human-readable text. The method differs from earlier interpretability tools…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-06-16/anthropic-s-autoencoders-translate-model-activations-into-re (subscription information: https://polylog.news/pricing).