# New Benchmarks Ask Whether Agents Actually Learn Anything Between Tasks

Three papers posted the same day target the same gap: evaluations reset the agent after every prompt, so nothing measures accumulated experience or the cost of memory.

- Published: 2026-09-09T06:20:57.202Z
- Canonical: https://polylog.news/ai/2026-09-09/new-benchmarks-ask-whether-agents-actually-learn-anything-be
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.LG](https://arxiv.org/abs/2609.05435), [arXiv cs.AI](https://arxiv.org/abs/2609.05441), [arXiv cs.CL](https://arxiv.org/abs/2609.05531), [Polylog editors](https://polylog.news)

Three papers posted to arXiv on September 9 converge on the same complaint about agent evaluation. AhaBench argues that agents are expected to work over long horizons, reusing worked examples and adapting to delayed consequences, while most…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-09-09/new-benchmarks-ask-whether-agents-actually-learn-anything-be (subscription information: https://polylog.news/pricing).