# Benchmarks Move From Tasks to Whole Jobs

Agent evaluation shifts from isolated task suites toward simulated occupations scored cumulatively over long sequences, and expect recurring job-shaped benchmarks that reveal a persistent gap between strong per-task scores and the ability to hold a role without degrading.

- Conviction: 40 / 100 (forming)
- Horizon: Emerging (watchlist)
- Tracking since: 2026-09-14T00:00:00.000Z
- Last updated: 2026-09-14T14:04:09.672Z
- Canonical: https://polylog.news/ai/trends/agent-job-role-benchmarks
- Publisher: Polylog
- Affected regions: Global

## Recent evidence

- [confirms] A New Benchmark Asks Whether an Agent Can Learn an Actual Job, Not Complete a Task (2026-09-14): ApprenticeBench drops a computer-use agent into an accounts payable role and scores cumulative success across 100 vendor bills processed in sequence, rather than scoring isolated tasks. Cumulative scoring over a sequence is the design choice that exposes degradation across a shift, which per-task suites structurally cannot measure.
