← Trends

Benchmarks Move From Tasks to Whole Jobs

Agent evaluation shifts from isolated task suites toward simulated occupations scored cumulatively over long sequences, and expect recurring job-shaped benchmarks that reveal a persistent gap between strong per-task scores and the ability to hold a role without degrading.

forming · confidence 40 · Emerging (watchlist) · tracking since September 14, 2026 · updated September 14, 2026

Sign in to get threshold and movement alerts for this trend.

Why the conviction moved

  • Sep 14
    Strengthened +5

    ApprenticeBench drops a computer-use agent into an accounts payable role and scores cumulative success across 100 vendor bills processed in sequence, rather than scoring isolated tasks. Cumulative scoring over a sequence is the design choice that exposes degradation across a shift, which per-task suites structurally cannot measure.

Source trail

  • Supporting · September 14, 2026

    A New Benchmark Asks Whether an Agent Can Learn an Actual Job, Not Complete a Task

    ApprenticeBench drops a computer-use agent into an accounts payable role and scores cumulative success across 100 vendor bills processed in sequence, rather than scoring isolated tasks. Cumulative scoring over a sequence is the design choice that exposes degradation across a shift, which per-task suites structurally cannot measure.

    AI ML Big Data (Telegram)

Unlock full source trail, score history, and daily updates.

Unlock Trends

Affected regions & assets

RegionsGlobal

Townsquare

Argue the thesis in Townsquare.