← Trends

The Harness Becomes the Capability Layer

An increasing share of measured agent capability comes from orchestration code rather than model weights, so expect harness-aware benchmarks, harness-specific pricing, and disputes over how much of any announced gain belongs to the model.

weakening · confidence 28 · Emerging (watchlist) · tracking since August 22, 2026 · updated August 28, 2026

Sign in to get threshold and movement alerts for this trend.

Score history

Daily conviction score, 0 to 100. Higher means the thesis is more strongly corroborated.

Aug 27 · 30Aug 28 · 28

Now 28 · -2 since Aug 27 · ranged 28 to 30

Showing the last few days. Unlock full score history.

Why the conviction moved

  • Aug 22
    Strengthened +8

    Nvidia reported clearing all 183 public ARC-AGI-3 levels by wrapping Claude Opus 5 in its own AVO harness, versus the 30.2 percent the ARC Prize Foundation independently measured for the same model. The gap is attributable entirely to orchestration code rather than weights, and because Nvidia both generated and scored the run it is exactly the harness-attribution dispute the thesis predicts.

  • Aug 22
    Weakened

    Scale AI's benchmark on models rewriting another agent's harness found across 111 scored runs that the choice of optimizer model separated results more than the coding harness it worked through, and that native harnesses were not consistently better than a shared one. That is direct measured evidence that weights, not orchestration code, still dominate agent capability.

Source trail

  • Supporting · August 22, 2026

    Nvidia Reports a Perfect Public-Set Score on ARC-AGI-3 by Wrapping Claude Opus 5 in Its Own Agent

    Nvidia reported clearing all 183 public ARC-AGI-3 levels by wrapping Claude Opus 5 in its own AVO harness, versus the 30.2 percent the ARC Prize Foundation independently measured for the same model. The gap is attributable entirely to orchestration code rather than weights, and because Nvidia both generated and scored the run it is exactly the harness-attribution dispute the thesis predicts.

    NVIDIA Technical Blog
  • Contradicting · August 22, 2026

    Scale AI Benchmark Measures Whether Models Can Rewrite Another Agent's Harness

    Scale AI's benchmark on models rewriting another agent's harness found across 111 scored runs that the choice of optimizer model separated results more than the coding harness it worked through, and that native harnesses were not consistently better than a shared one. That is direct measured evidence that weights, not orchestration code, still dominate agent capability.

    arXiv

Unlock full source trail, score history, and daily updates.

Unlock Trends

Affected regions & assets

RegionsGlobal

Townsquare

Argue the thesis in Townsquare.