# Researchers Propose an Executable Benchmark for the Decisions Agents Make Before They Answer

Two arXiv papers target the choices an agent makes before answering (answer directly, decompose, retrieve, run code, delegate, verify or recover), which drive most of its cost.

- Published: 2026-08-04T06:16:46.804Z
- Canonical: https://polylog.news/ai/2026-08-04/researchers-propose-an-executable-benchmark-for-the-decision
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.LG](https://arxiv.org/abs/2608.00106), [arXiv cs.LG](https://arxiv.org/abs/2608.00107)

Most agent benchmarks score the final answer. Two papers posted on August 4 argue that the more consequential behavior is what the controller decides to do first. Learning Compositional Meta-Routing for Agentic Workflows frames the problem…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-08-04/researchers-propose-an-executable-benchmark-for-the-decision (subscription information: https://polylog.news/pricing).