# Scale AI Benchmark Measures Whether Models Can Rewrite Another Agent's Harness

Across 111 scored runs, the choice of optimizer model separated results more than the coding harness the optimizer worked through, and native harnesses were not consistently better than a shared one.

- Published: 2026-08-22T06:18:28.650Z
- Canonical: https://polylog.news/ai/2026-08-22/scale-ai-benchmark-measures-whether-models-can-rewrite-anoth
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv](https://arxiv.org/abs/2608.06301), [Polylog editors](https://polylog.news)

Scale AI published HarnessOpt-Bench, a benchmark for automated harness optimization. In the setup, an optimizer, defined as a large language model paired with a coding harness, is given a target agent's starting harness, graded feedback on…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-08-22/scale-ai-benchmark-measures-whether-models-can-rewrite-anoth (subscription information: https://polylog.news/pricing).