Nvidia just showed that the harness, not the AI model, is now the real hero
Summarized from techcrunch.com
Nvidia’s recent research suggests that the harness, or software wrapper around an AI model, is more important than the underlying model itself when it comes to achieving long-horizon tasks. The harness includes tools, memory management, and rules that transform a raw model into something capable of acting independently. By using a custom harness tweaked for effective memory handling and incorporating a “supervisor” component, researchers were able to get Claude Opus 5 to achieve a perfect score on the interactive reasoning benchmark ARC-AGI-3, which involves figuring out how to play and win in 2D games without instructions.
The research highlights that while model choice does matter, it is not the only factor determining an AI’s performance, especially for long-horizon tasks. The harness plays a crucial role by handling memory, context, and feedback, essentially making a model into an agent. OpenAI also conducted its own research on long-horizon tasks and found that simply tweaking two settings in the harness could triple their models’ scores, though none of them reached the 100% score achieved by Nvidia’s researchers.
The concept of a supervising agent is not entirely new but has been underutilized. A supervisor can nudge an agent when it gets stuck or starts exploring paths leading to dead ends, acting almost like a CEO would. While this idea isn’t novel, today most users rely on only one layer for their harness, such as Claude Code or Hermes. Nvidia’s Agentic Variation Operators (AVO) is an example of a more advanced and customized harness designed specifically for agentic tasks. This research underscores that model choice is just one part of the equation; the harness also plays a critical role in determining the effectiveness of AI agents, particularly in long-horizon tasks where stringing decisions together over time is essential.