# A Wave of Benchmarks Probes How Easily AI Agents Are Manipulated

New work targets code-review agents, deceptive shopping interfaces, and streaming guardrails, alongside a real incident where an unsupervised agent ran up a large cloud bill.

- Published: 2026-06-15T07:00:34.492Z
- Canonical: https://polylog.news/ai/2026-06-15/a-wave-of-benchmarks-probes-how-easily-ai-agents-are-manipul
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.CR](https://arxiv.org/abs/2606.13757), [arXiv cs.CL](https://arxiv.org/abs/2606.13686), [Polylog editors](https://polylog.news)

Several papers posted the same day converge on a single theme, that autonomous agents fail under adversarial pressure in ways static benchmarks miss. SEVRA-BENCH studies the social engineering of large language model (LLM) reviewers used in…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-06-15/a-wave-of-benchmarks-probes-how-easily-ai-agents-are-manipul (subscription information: https://polylog.news/pricing).