Nvidia’s AVO Harness Pushes Claude Opus 5 to a Perfect Score on ARC-AGI-3
Nvidia's new AVO agent architecture completed every level of the ARC-AGI-3 benchmark using Claude Opus 5, a model that scores roughly 30 percent on the same test without it.
Nvidia researchers have built an agent architecture that takes Claude Opus 5, a model that scores around 30 percent on its own, and pushes it to a perfect 100 percent on ARC-AGI-3, one of the hardest benchmarks yet devised for autonomous AI agents. The result, published on Nvidia’s developer blog on August 21, 2026, is Nvidia’s clearest public demonstration that the software wrapped around a language model, not the model itself, decides whether an AI agent can handle long, open-ended tasks.
Table Of Content
What ARC-AGI-3 actually tests
ARC-AGI-3 is an interactive reasoning benchmark built by the ARC Prize Foundation. Earlier ARC-AGI tests were static: an agent saw a handful of input and output grid examples and had to produce the matching output for a new grid. ARC-AGI-3 works differently. It drops an agent into a turn-based game environment with no instructions, no stated rules, and no declared goal. The agent has to explore on its own, work out how the environment behaves, figure out what winning even looks like, and carry that understanding forward as the levels get harder. Agents are scored on how many levels they finish and, separately, on how efficiently they get there.
The harness, not the model, did the work
Nvidia’s system is called AVO, short for Agentic Variation Operators. It is not a new model. It is an architecture built around an existing one, in this case Anthropic’s Claude Opus 5. AVO runs an iterative loop: inspect the current state, plan the next move, implement it, then evaluate the outcome before cycling again. It keeps persistent memory across that loop, carrying forward prior attempts, evaluation results, and accumulated reasoning instead of starting fresh on every turn. A separate supervisor component watches the broader trajectory for stagnation or repeated unproductive cycles and can redirect the main agent toward a different strategy when needed.
On ARC-AGI-3, AVO completed all 183 levels across the benchmark’s 25 public environments, reaching a perfect 100.00 on the benchmark’s Relative Human Action Efficiency score, a measure of how many actions an agent needs to clear each level relative to a first-time human baseline. It did so in 6,624 environment actions, about 12 percent fewer than the 7,542 actions used by VISTA, a rival harness system that also cleared every level on the benchmark but needed more actions to do it. Claude Opus 5 on its own, without AVO’s scaffolding, scores about 30 percent on ARC-AGI-3 at its highest reasoning-effort setting, according to Nvidia’s blog post, which cites the ARC Prize’s own published figures for the model. TechCrunch reported that this unassisted 30 percent was already the best score among all the standalone models tested on the benchmark, before any harness was added.
Built for chip design, tested on a video game
AVO was not built for game-playing benchmarks. Nvidia says the system was first developed for software engineering and GPU-kernel optimization work: tasks where an agent has to inspect existing code, form a hypothesis, make a change, run hardware-grounded tests, interpret the results, and repeatedly revise its approach. In one such test, an attention-kernel study, Nvidia says AVO ran continuously for seven days, explored more than 500 optimization directions, and produced 40 committed kernel versions. On Nvidia’s own DGX B200 systems, the resulting kernels outperformed cuDNN by up to 3.5 percent and FlashAttention-4 by up to 10.5 percent.
Nvidia is explicit that AVO itself is a research project, not a shipping product. The company instead sells pieces of the underlying tooling under its Nemo brand, some of it commercial and much of it openly available. “The model matters, but the model is not the entire agent,” Nvidia researchers Terry Chen, Yeyin (Eva) Zhu, Zhifan Ye, Jean-Francois Puget, and Humphrey Shi wrote in the blog post announcing the result.
Why the industry keeps circling back to the harness
Nvidia is not alone in reaching that conclusion. Speaking to TechCrunch, Adel El Hallak, vice president of product in Nvidia’s AI unit, argued that most people misjudge what an agent actually is. “Generally speaking, the world interprets an agent almost as an API of the model,” he said. He added that an agent is more than that: “It is the model. It is the scaffolding around the model, which we call the harness, i.e. the set of tools that it utilizes. It is the runtime and the associated skills and libraries that we give it access to.”
TechCrunch also reported that OpenAI, whose own models score less than 10 percent on ARC-AGI-3 unassisted, ran a similar investigation last month and found that adjusting just two harness settings tripled its models’ scores, though still well short of AVO’s perfect result. Databricks has reached a related conclusion from the cost side: its CEO, Ali Ghodsi, told TechCrunch that harness choice alone can double what it costs to run the same model in production. “You can pick the same model but different harnesses, and you get significantly more cost if you use the wrong harness,” he said. “So you think, oh, this is an expensive model. This is a cheap model. But wait, which harness are you using? That itself can 2x your cost.”
Nvidia’s pitch is that open, swappable harnesses give developers more control than one bundled tightly with a single model provider, an argument that also serves as a pitch for its own Nemo tooling. But the ARC-AGI-3 result gives that argument something concrete to point to: Claude Opus 5’s bare 30 percent was already the best unassisted score Nvidia and TechCrunch cited among the models tested, and AVO’s harness still turned it into a perfect one.








No Comment! Be the first one.