Morning Edition · Saturday, August 22, 2026Published at 2:18 AM EDT · New York
The same model the ARC Prize Foundation independently measured at 30.2 percent cleared all 183 public levels once Nvidia wrapped it in its AVO harness, a result Nvidia generated and scored on its own.

Nvidia said on August 21 that AVO, a general-purpose agent built around what the company calls Agentic Variation Operators, completed all 183 levels across the 25 public environments of ARC-AGI-3, the interactive reasoning benchmark from the Abstraction and Reasoning Corpus for Artificial General Intelligence (ARC-AGI) family. AVO scored 100.00 on the benchmark's Relative Human Action Efficiency metric and, according to Nvidia, needed 12 percent fewer actions than VISTA, the previous best-performing system. The benchmark places an agent in game-like environments with no instructions, no stated rules and no declared goal, so the agent has to work out for itself what its actions do.
The detail that matters for engineers is what the harness added. Nvidia built AVO around Anthropic's Claude Opus 5, a model the ARC Prize Foundation itself measured at 30.2 percent on the same benchmark. AVO was originally designed for autonomous software engineering and graphics processing unit (GPU) kernel optimization, and the reported gap between the bare model and the wrapped model is roughly a factor of three.
Two qualifications apply to that number. Nvidia ran the evaluation on its own reimplementation of the task interface and states plainly that these results do not cover the semi-private or fully private competition sets. The ARC Prize testing policy permits this kind of self-reported public-set run, provided the reporter discloses it, and reserves official leaderboard status for runs the foundation administers itself. The public environments exist specifically so outside researchers can study them, which is also why the private sets exist.
The claim spread quickly through developer channels, including a widely shared summary of Nvidia's announcement that largely omitted the public-set caveat.
Start a discussion in Townsquare.
More from this edition
Nvidia, which sells the accelerators and now an agent architecture on top of them, gains if buyers conclude that harness engineering, not model weights, sets frontier capability, and the 30.2 percent to 100 percent gap transfers value from model vendors to the vendor of the scaffold.
The number is real and Nvidia disclosed its limits, but the run used Nvidia's own reimplementation of the task interface on the public set, and the ARC Prize Foundation does not independently verify self-reported scores, so the comparison to the foundation-administered 30.2 percent figure is not like for like.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
If a scaffold can triple a frontier model's score on an interactive benchmark, the thing being measured is the agent system, not the underlying model weights, and that raises the value of harness engineering relative to the value of a model subscription. Nvidia gains from that framing directly, since it sells both the computing hardware and now the agent architecture built on top of it. Model vendors lose pricing power if buyers conclude that a mid-tier model paired with a strong harness can match a premium model paired with a weak one. Anthropic is exposed on both sides of this: its model did the reasoning, and another company received the public credit for the result.
What to watch
Observations to monitor, not financial advice.
Synthesized from: NVIDIA Technical Blog · The New Stack · Polylog editors · ARC Prize verified testing policy · Anthropic
Comments
1Aug 22, 9:06 AM · edited
The jump from 30.2 percent without AVO to 100 percent with it was measured on the 183 public levels that are fixed and known in advance, and Nvidia has not published a score on the private held out set.