Nvidia research finds AI harness can matter more than model choice

Via nvidia.com

Nvidia research finds AI harness can matter more than model choice

A custom wrapper with memory controls and a supervisory agent lifted Claude Opus 5’s ARC-AGI-3 score from 30% to 100%

Nvidia research suggests the software harness around an AI model can matter more than the model itself when agents handle long-horizon tasks. A harness supplies the tools, memory management and rules that let a model act autonomously.

Researchers used a custom harness with memory controls and a supervisory component to help Claude Opus 5 complete ARC-AGI-3, an interactive reasoning benchmark built from 2D games with no instructions. The model scored 100% with the harness, compared with 30% without it.

Advertisement

The findings indicate that model choice is only one part of agent performance. The harness manages memory, context and feedback, while the supervisory agent can redirect the system when it gets stuck or starts down an unproductive path.

Nvidia vice president Adel El Hallack described the harness as the scaffolding around the model, including its tools, runtime, skills and libraries. Nvidia researchers called their system Agentic Variation Operators, or AVO.

The research is not a new Nvidia product. The company provides open and commercial components for building harnesses through its NeMo ecosystem. 

Other research, including work from OpenAI and Databricks, has also found that harness design can significantly affect agent accuracy and cost.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.
Nvidia research finds AI harness can matter more than model choice
Nvidia research finds AI harness can matter more than model choice

A custom wrapper with memory controls and a supervisory agent lifted Claude Opus 5’s ARC-AGI-3 score from 30% to 100%

Share

Add us on Google

Via nvidia.com

Nvidia research suggests the software harness around an AI model can matter more than the model itself when agents handle long-horizon tasks. A harness supplies the tools, memory management and rules that let a model act autonomously.

Researchers used a custom harness with memory controls and a supervisory component to help Claude Opus 5 complete ARC-AGI-3, an interactive reasoning benchmark built from 2D games with no instructions. The model scored 100% with the harness, compared with 30% without it.

Advertisement

The findings indicate that model choice is only one part of agent performance. The harness manages memory, context and feedback, while the supervisory agent can redirect the system when it gets stuck or starts down an unproductive path.

Nvidia vice president Adel El Hallack described the harness as the scaffolding around the model, including its tools, runtime, skills and libraries. Nvidia researchers called their system Agentic Variation Operators, or AVO.

The research is not a new Nvidia product. The company provides open and commercial components for building harnesses through its NeMo ecosystem. 

Other research, including work from OpenAI and Databricks, has also found that harness design can significantly affect agent accuracy and cost.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.