TRENDING
Rows of identical brass-colored apartment mailboxes with small locks and name labels along an orange corridor wall
October 9, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
Street-level upward view of the Monetary Authority of Singapore building and neighbouring office towers under a pale sky
October 9, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
Cast-iron late Qing dynasty coin minting press with a large flywheel, displayed in a museum case
October 9, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google
Rows of closed oak library card catalog drawers, each with a brass pull and a blank label holder
October 9, 2026
How to Encrypt PII in Python and Keep It Searchable With Blind Indexes
Close-up of a vintage Western Electric manual telephone switchboard with orange lamps, red patch cords plugged into jacks, a rotary dial and a black handset
October 9, 2026
Microsoft’s Agent Lightning v1.0 Turns Agent Training Into a Sample-Accounting Problem
09 Oct 2026
SXZ.io SXZ.io
  • Home
Search the Site
Popular Searches:
Technology Amazon AI
Recent Posts
Two orange safety relief valves on grey pressure vessels in an industrial plant
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
Yellow diamond-shaped merging traffic warning sign showing a side road joining a main road
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A lugworm lying on wet sand and mud at low tide
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
SXZ.io SXZ.io
  • Home

Categories

Articles 232 Posts
News 234 Posts
Learning Hub 204 Posts
Home/News/Nvidia’s AVO Harness Pushes Claude Opus 5 to a Perfect Score on ARC-AGI-3
News

Nvidia’s AVO Harness Pushes Claude Opus 5 to a Perfect Score on ARC-AGI-3

Nvidia's new AVO agent architecture completed every level of the ARC-AGI-3 benchmark using Claude Opus 5, a model that scores roughly 30 percent on the same test without it.

August 22, 2026 4 Min Read
44

Nvidia researchers have built an agent architecture that takes Claude Opus 5, a model that scores around 30 percent on its own, and pushes it to a perfect 100 percent on ARC-AGI-3, one of the hardest benchmarks yet devised for autonomous AI agents. The result, published on Nvidia’s developer blog on August 21, 2026, is Nvidia’s clearest public demonstration that the software wrapped around a language model, not the model itself, decides whether an AI agent can handle long, open-ended tasks.

Table Of Content

  • What ARC-AGI-3 actually tests
  • The harness, not the model, did the work
  • Built for chip design, tested on a video game
  • Why the industry keeps circling back to the harness

What ARC-AGI-3 actually tests

ARC-AGI-3 is an interactive reasoning benchmark built by the ARC Prize Foundation. Earlier ARC-AGI tests were static: an agent saw a handful of input and output grid examples and had to produce the matching output for a new grid. ARC-AGI-3 works differently. It drops an agent into a turn-based game environment with no instructions, no stated rules, and no declared goal. The agent has to explore on its own, work out how the environment behaves, figure out what winning even looks like, and carry that understanding forward as the levels get harder. Agents are scored on how many levels they finish and, separately, on how efficiently they get there.

The harness, not the model, did the work

Nvidia’s system is called AVO, short for Agentic Variation Operators. It is not a new model. It is an architecture built around an existing one, in this case Anthropic’s Claude Opus 5. AVO runs an iterative loop: inspect the current state, plan the next move, implement it, then evaluate the outcome before cycling again. It keeps persistent memory across that loop, carrying forward prior attempts, evaluation results, and accumulated reasoning instead of starting fresh on every turn. A separate supervisor component watches the broader trajectory for stagnation or repeated unproductive cycles and can redirect the main agent toward a different strategy when needed.

On ARC-AGI-3, AVO completed all 183 levels across the benchmark’s 25 public environments, reaching a perfect 100.00 on the benchmark’s Relative Human Action Efficiency score, a measure of how many actions an agent needs to clear each level relative to a first-time human baseline. It did so in 6,624 environment actions, about 12 percent fewer than the 7,542 actions used by VISTA, a rival harness system that also cleared every level on the benchmark but needed more actions to do it. Claude Opus 5 on its own, without AVO’s scaffolding, scores about 30 percent on ARC-AGI-3 at its highest reasoning-effort setting, according to Nvidia’s blog post, which cites the ARC Prize’s own published figures for the model. TechCrunch reported that this unassisted 30 percent was already the best score among all the standalone models tested on the benchmark, before any harness was added.

Built for chip design, tested on a video game

AVO was not built for game-playing benchmarks. Nvidia says the system was first developed for software engineering and GPU-kernel optimization work: tasks where an agent has to inspect existing code, form a hypothesis, make a change, run hardware-grounded tests, interpret the results, and repeatedly revise its approach. In one such test, an attention-kernel study, Nvidia says AVO ran continuously for seven days, explored more than 500 optimization directions, and produced 40 committed kernel versions. On Nvidia’s own DGX B200 systems, the resulting kernels outperformed cuDNN by up to 3.5 percent and FlashAttention-4 by up to 10.5 percent.

Nvidia is explicit that AVO itself is a research project, not a shipping product. The company instead sells pieces of the underlying tooling under its Nemo brand, some of it commercial and much of it openly available. “The model matters, but the model is not the entire agent,” Nvidia researchers Terry Chen, Yeyin (Eva) Zhu, Zhifan Ye, Jean-Francois Puget, and Humphrey Shi wrote in the blog post announcing the result.

Why the industry keeps circling back to the harness

Nvidia is not alone in reaching that conclusion. Speaking to TechCrunch, Adel El Hallak, vice president of product in Nvidia’s AI unit, argued that most people misjudge what an agent actually is. “Generally speaking, the world interprets an agent almost as an API of the model,” he said. He added that an agent is more than that: “It is the model. It is the scaffolding around the model, which we call the harness, i.e. the set of tools that it utilizes. It is the runtime and the associated skills and libraries that we give it access to.”

TechCrunch also reported that OpenAI, whose own models score less than 10 percent on ARC-AGI-3 unassisted, ran a similar investigation last month and found that adjusting just two harness settings tripled its models’ scores, though still well short of AVO’s perfect result. Databricks has reached a related conclusion from the cost side: its CEO, Ali Ghodsi, told TechCrunch that harness choice alone can double what it costs to run the same model in production. “You can pick the same model but different harnesses, and you get significantly more cost if you use the wrong harness,” he said. “So you think, oh, this is an expensive model. This is a cheap model. But wait, which harness are you using? That itself can 2x your cost.”

Nvidia’s pitch is that open, swappable harnesses give developers more control than one bundled tightly with a single model provider, an argument that also serves as a pitch for its own Nemo tooling. But the ARC-AGI-3 result gives that argument something concrete to point to: Claude Opus 5’s bare 30 percent was already the best unassisted score Nvidia and TechCrunch cited among the models tested, and AVO’s harness still turned it into a perfect one.

Tags:

Agentic AIAI AgentsAnthropicBenchmarksNVIDIA

Share

A drum seismograph recording a calm baseline pattern suddenly interrupted by a burst of erratic spikes
Previous Post

How to Turn Slow SQL Queries Into Actionable Reliability Metrics With OpenTelemetry

Close-up profile of the wooden Trojan Horse replica at the Canakkale waterfront in Turkey, showing weathered planks and rope bindings against a blue sky
Next Post

Adversa AI’s Cryptographic Context Injection Turns Encryption Into a Guardrail Blind Spot

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest
08 Oct
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
08 Oct
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
Trending
October 8, 2026
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
October 8, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
October 8, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google

Related Posts

Rows of server racks in a data center representing network infrastructure targeted by botnets
News

C0XMO Botnet Shows Why Old Router Firmware Still Matters

June 7, 2026
Close-up of a USB flash drive, representing physical data-theft risk in office security incidents
News

Fake IT Support Is Now Walking Through the Front Door

June 7, 2026
A phone security app on a smartphone resting on a laptop keyboard.
News

Everest Forms Pro Flaw Is Being Exploited to Create Rogue WordPress Admins

June 7, 2026
A customer-support representative wearing a headset against a dark studio background.
Articles

The Meta AI Support Hack Was a Plain Old Authorization Failure

June 7, 2026
SXZ.io SXZ.io
  • [email protected]

Categories

Articles
Learning Hub
News

All Rights Reserved by SXZ.io ©2026