# Researchers Report That Agent Models Internally Register When They Have Been Prompt-Injected

A preprint argues the hidden states of agentic language models carry a detectable signal of exposure to malicious instructions hidden in tool output, even when the agent goes on to obey them.

- Published: 2026-08-05T05:32:18.568Z
- Canonical: https://polylog.news/ai/2026-08-05/researchers-report-that-agent-models-internally-register-whe
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv](https://arxiv.org/abs/2608.02657), [Anthropic](https://www.anthropic.com/research/team/frontier-red-team)

Indirect prompt injection remains the unsolved failure mode of production agents. An agent reads a web page, a document or a tool response, that content contains instructions, and the model treats them as if the user had written them. Every…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-08-05/researchers-report-that-agent-models-internally-register-whe (subscription information: https://polylog.news/pricing).