# Three Preprints Argue Agent Evaluation Is Measuring the Wrong Thing

Outcome-only judging cannot see an agent that reaches the right answer by an unacceptable route, and graphical interface world models are tested one step at a time while being used as multi-step environments.

- Published: 2026-09-02T06:20:39.906Z
- Canonical: https://polylog.news/ai/2026-09-02/three-preprints-argue-agent-evaluation-is-measuring-the-wron
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.CL](https://arxiv.org/abs/2609.00038), [arXiv cs.CL](https://arxiv.org/abs/2609.00048), [arXiv cs.AI](https://arxiv.org/abs/2609.00002)

Three preprints posted the same day converge on a single complaint about how the industry measures agents. The first targets the production default. Outcome-only evaluation shows a judge the request and the final reply and asks whether the…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-09-02/three-preprints-argue-agent-evaluation-is-measuring-the-wron (subscription information: https://polylog.news/pricing).