The Harness Becomes the Capability Layer
An increasing share of measured agent capability comes from orchestration code rather than model weights, so expect harness-aware benchmarks, harness-specific pricing, and disputes over how much of any announced gain belongs to the model.
weakening · confidence 28 · Emerging (watchlist) · tracking since August 22, 2026 · updated August 28, 2026
Score history
Daily conviction score, 0 to 100. Higher means the thesis is more strongly corroborated.
Now 28 · -2 since Aug 27 · ranged 28 to 30
Showing the last few days. Unlock full score history.
Why the conviction moved
- Aug 22Strengthened +8
Nvidia reported clearing all 183 public ARC-AGI-3 levels by wrapping Claude Opus 5 in its own AVO harness, versus the 30.2 percent the ARC Prize Foundation independently measured for the same model. The gap is attributable entirely to orchestration code rather than weights, and because Nvidia both generated and scored the run it is exactly the harness-attribution dispute the thesis predicts.
- Aug 22Weakened
Scale AI's benchmark on models rewriting another agent's harness found across 111 scored runs that the choice of optimizer model separated results more than the coding harness it worked through, and that native harnesses were not consistently better than a shared one. That is direct measured evidence that weights, not orchestration code, still dominate agent capability.
Source trail
Supporting · August 22, 2026
Nvidia Reports a Perfect Public-Set Score on ARC-AGI-3 by Wrapping Claude Opus 5 in Its Own Agent
Nvidia reported clearing all 183 public ARC-AGI-3 levels by wrapping Claude Opus 5 in its own AVO harness, versus the 30.2 percent the ARC Prize Foundation independently measured for the same model. The gap is attributable entirely to orchestration code rather than weights, and because Nvidia both generated and scored the run it is exactly the harness-attribution dispute the thesis predicts.
NVIDIA Technical BlogContradicting · August 22, 2026
Scale AI Benchmark Measures Whether Models Can Rewrite Another Agent's Harness
Scale AI's benchmark on models rewriting another agent's harness found across 111 scored runs that the choice of optimizer model separated results more than the coding harness it worked through, and that native harnesses were not consistently better than a shared one. That is direct measured evidence that weights, not orchestration code, still dominate agent capability.
arXiv
Unlock full source trail, score history, and daily updates.
Unlock TrendsAffected regions & assets
Townsquare
Argue the thesis in Townsquare.