# Paper Finds Agent Models Internally Encode When Their Tool Output Was Poisoned

The authors report that hidden activations in agentic language models carry a detectable signal of indirect prompt-injection exposure, suggesting a cheap runtime monitor rather than another input filter.

- Published: 2026-08-05T05:48:38.661Z
- Canonical: https://polylog.news/ai/2026-08-05/paper-finds-agent-models-internally-encode-when-their-tool-o
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.CR](https://arxiv.org/abs/2608.02657</source_url_placeholder), [arXiv](https://arxiv.org/abs/2608.02657)

A preprint posted to arXiv, Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure, takes a different angle on the most persistent security problem in agent deployment. Indirect prompt injection places a mali…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-08-05/paper-finds-agent-models-internally-encode-when-their-tool-o (subscription information: https://polylog.news/pricing).