# New Benchmarks Target the Decision Agents Make Before They Answer

Two arXiv releases score the meta-decision of whether to answer, decompose, retrieve, execute code or delegate. A third replaces static persona prompts with synthesized lifelong memory.

- Published: 2026-08-04T06:03:13.954Z
- Canonical: https://polylog.news/ai/2026-08-04/new-benchmarks-target-the-decision-agents-make-before-they-a
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.LG](https://arxiv.org/abs/2608.00107), [arXiv cs.LG](https://arxiv.org/abs/2608.00106), [arXiv cs.CL](https://arxiv.org/abs/2608.00007)

Agentic systems spend most of their token budget on choices that precede the answer. MetaRoute-Bench formalizes exactly those choices, evaluating whether a controller should answer directly, decompose a task, call a tool, run code, delegate…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-08-04/new-benchmarks-target-the-decision-agents-make-before-they-a (subscription information: https://polylog.news/pricing).