# New Benchmark Tests Whether Computer-Use Agents Refuse Traps and Ask When Instructions Are Vague

ADeptS-Bench pairs benign and malicious versions of the same task with threats embedded in the visual interface, and separately scores whether an agent seeks clarification instead of guessing.

- Published: 2026-08-28T06:11:21.137Z
- Canonical: https://polylog.news/ai/2026-08-28/new-benchmark-tests-whether-computer-use-agents-refuse-traps
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.CR](https://arxiv.org/abs/2608.26204)

Agents that drive real desktop and mobile applications on a user's behalf are now shipping in commercial products, and the evaluations covering them mostly measure task completion. ADeptS-Bench, posted to arXiv on Friday, tests two things t…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-08-28/new-benchmark-tests-whether-computer-use-agents-refuse-traps (subscription information: https://polylog.news/pricing).