# New Benchmarks Try to Score Agents on Execution Rather Than Multiple Choice

One framework runs coding agents against live backtests in algorithmic trading, another replaces roleplay scores with per-requirement checklists tied to dialogue evidence.

- Published: 2026-08-14T06:27:18.545Z
- Canonical: https://polylog.news/ai/2026-08-14/new-benchmarks-try-to-score-agents-on-execution-rather-than
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.CL](https://arxiv.org/abs/2608.11232), [arXiv cs.CL](https://arxiv.org/abs/2608.11236)

Two evaluation papers this week share a diagnosis: static benchmarks for agents are contaminated by training data, and a single aggregate score tells a deploying team nothing about which requirement failed. Backtrader-Bench addresses both i…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-08-14/new-benchmarks-try-to-score-agents-on-execution-rather-than (subscription information: https://polylog.news/pricing).