NVIDIA's AVO agent architecture clears all 25 ARC-AGI-3 public-set environments with a 100.00 RHAE score and all 183 levels
Related research and updatesSynopsis
NVIDIA transferred its general-purpose long-horizon agent architecture AVO from GPU-kernel optimization to the ARC-AGI-3 interactive reasoning benchmark, completing all 25 public-set environments and all 183 levels with a 100.00 RHAE score in 6,624 environment actions, and reports that persistent memory and supervision sustain autonomous progress across domains.
Interpretation
The same AVO architecture sustained long-horizon autonomous work on two very different tasks: in the attention-kernel study it ran continuously for seven days, explored more than 500 optimization directions, and produced 40 committed kernel versions; on the ARC-AGI-3 public set it completed all 25 environments and all 183 levels. Prior long-horizon agent work is often tied to a specific domain; this work places GPU-kernel optimization and interactive reasoning under one architecture, indicating that what transfers is the machinery for sustained autonomous progress rather than domain knowledge. Grounded in scale metrics from two full runs: seven days of continuous operation, more than 500 explored directions, and 40 committed kernel versions, plus completion of 25 environments and 183 levels on the public set; the authors state this is not a controlled ablation.
AVO scored 100.00 RHAE on the ARC-AGI-3 public set, solving all 183 levels in 6,624 environment actions; the cited VISTA reports 7,542 environment actions on the same public-set levels, so AVO used approximately 12% fewer actions. It separates model evaluation from agent-system evaluation: the same model family performs differently under different agent systems and evaluation setups, with ARC Prize separately reporting approximately 30% for Claude Opus 5 at High reasoning effort. A cross-system comparison; the authors explicitly note the two systems differ in agent backend, observation representation, memory, context management, and other implementation details, so it is not a controlled ablation and the 12% difference cannot be attributed to a single component.
AVO sustains progress beyond a single context through two mechanisms: persistent memory carries forward prior implementations, evaluation results, compiler and profiler outputs, and accumulated reasoning; a supervisor monitors the broader trajectory for stagnation or repeated unproductive cycles and redirects the main agent. It frames long-horizon capability as a system property: memory determines what survives, tools determine what actions are possible, feedback grounds progress, and recovery lets work continue beyond a single model invocation. Mechanism descriptions paired with role division during the seven-day kernel run (the main agent decided what to inspect, change, test, and evaluate, while the supervisor helped maintain forward progress when the search plateaued); the text notes the experiment does not isolate the memory system's contribution.
In the ARC-AGI-3 configuration the model operated in a text-only modality, receiving each observation as an exact 64 x 64 text grid with no images or image tokens, and received available actions without descriptions of rules or goals, having to infer their effects through interaction. Unlike VISTA's primary configuration using a rendered 512 x 512 PNG, and unlike Tycho's explicit executable world-model route, this work adopts direct-interaction design principles and independently reimplements the task interface. Explicit description of the interface and modality, plus contrast with the VISTA and Tycho routes; the authors note several task-interface elements were informed by VISTA while the agent backend is fundamentally different.
Perspective
The result applies to the 25-environment ARC-AGI-3 public set using the official scorecard and RHAE metric, and does not cover the semi-private or fully private competition sets. For readers, it offers a reference case for reusing one agent architecture on long-horizon, feedback-driven tasks: settings that need persistent state, tool interfaces, execution-grounded validation, and failure recovery, such as software engineering and GPU-kernel optimization, as well as interactive environments where rules and goals are not stated. The text also treats cross-model operation as a design goal and reports preliminary paired observations with GPT-5.6 Sol on a subset of games for later system comparisons.
The cross-system action comparison (6,624 versus 7,542) involves differences in agent backend, observation representation, memory, and context management; the authors explicitly state it is not a controlled ablation, so the 12% action difference cannot be attributed to one component. The approximately 30% reported by ARC Prize comes from the same model family under a different reasoning setting and a substantially different agent system and evaluation setup, so it is not a direct measure of AVO's contribution. The role of the memory system over long horizons is not isolated, and the paired experiments with GPT-5.6 Sol cover only a subset of games, leaving a broader systematic comparison to future work. In addition, this is a project introduction rather than a full experimental report, without complete tables or statistical detail; readers seeking reproduction should consult the cited paper and benchmark scoring methodology.
