# The Harness Becomes the Capability Layer

An increasing share of measured agent capability comes from orchestration code rather than model weights, so expect harness-aware benchmarks, harness-specific pricing, and disputes over how much of any announced gain belongs to the model.

- Conviction: 28 / 100 (weakening)
- Horizon: Emerging (watchlist)
- Tracking since: 2026-08-22T00:00:00.000Z
- Last updated: 2026-08-28T06:28:38.436Z
- Canonical: https://polylog.news/ai/trends/agent-harness-capability-layer
- Publisher: Polylog
- Affected regions: Global

## Recent score history

- 2026-08-27: 30
- 2026-08-28: 28

## Recent evidence

- [confirms] Nvidia Reports a Perfect Public-Set Score on ARC-AGI-3 by Wrapping Claude Opus 5 in Its Own Agent (2026-08-22): Nvidia reported clearing all 183 public ARC-AGI-3 levels by wrapping Claude Opus 5 in its own AVO harness, versus the 30.2 percent the ARC Prize Foundation independently measured for the same model. The gap is attributable entirely to orchestration code rather than weights, and because Nvidia both generated and scored the run it is exactly the harness-attribution dispute the thesis predicts.
- [contradicts] Scale AI Benchmark Measures Whether Models Can Rewrite Another Agent's Harness (2026-08-22): Scale AI's benchmark on models rewriting another agent's harness found across 111 scored runs that the choice of optimizer model separated results more than the coding harness it worked through, and that native harnesses were not consistently better than a shared one. That is direct measured evidence that weights, not orchestration code, still dominate agent capability.
