Skip to main content
Back to timeline
arXivSource publication:

VISTA visual harness lifts Claude Opus 5.0 on ARC-AGI-3 from a 40.68 to a perfect 100.00 Relative Human Action Efficiency and clears all 25 public games with 57.4% fewer actions

Related research and updates

Synopsis

The authors introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision by letting it perceive the environment through visual observations, keep a lossless visual memory of past observations in their original form, and actively retrieve and reorganize its visual input while reasoning; on ARC-AGI-3 it raises Claude Opus 5.0's Relative Human Action Efficiency from 40.68 to a perfect 100.00, completing all 25 public games with 57.4% fewer actions than first-time human participants, and it substantially outperforms baselines using the same underlying model with minimal harnesses across three additional benchmarks covering diverse visual games and puzzles.

Source-provided article image: VISTA: A Visual Harness for Reasoning in an Interactive World
Figure 1 ·

Figure 1: Comparison of different agent designs. (a) Language agents reason over textual observations. (b) Multimodal agents encode images into visual representations only once, but these representations may be compressed, lossy, and insufficient for subsequent reasoning, potentially limiting the models’ reasoning capabilities. (c) VISTA allows the multimodal model to iteratively call visual tools to retrieve past visual observations and reorganize its visual input as it reasons. For simplicity, we show the VLM as a vision transformer (ViT) encoder followed by an LLM. The two VLM blocks represent the same underlying model.

arXiv

Interpretation

VISTA gives a general-purpose multimodal model long-horizon vision: the model directly perceives the environment through visual observations, maintains a lossless visual memory that preserves past observations in their original form, and actively retrieves those observations and reorganizes its visual input as it reasons. Relative to prior interactive-agent approaches that rely on minimal harnesses or compressed memory, this work makes preserving raw visual observations plus active retrieval and reorganization the core harness mechanism, rather than having the model handle only the current frame or a lossy summary of history. The abstract supports this with both the mechanism description and quantitative ARC-AGI-3 results: Relative Human Action Efficiency rising from 40.68 to 100.00, all 25 public games completed, and 57.4% fewer actions than first-time human participants.

On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency from 40.68 to a perfect 100.00 and completes all 25 public games. This result directly quantifies the gap for the same underlying model with and without the VISTA harness, indicating the gain comes from the harness rather than a model change. The abstract reports specific scores (40.68 to 100.00), the number of games (25 public games), and the action savings (57.4%), forming a controlled comparison on the same benchmark.

VISTA's simple design extends naturally to diverse visual environments with minimal adaptation; across three additional benchmarks covering a diverse range of visual games and puzzles, it substantially outperforms baselines using the same underlying model with minimal harnesses. This indicates the harness is not a benchmark-specific solution but a general component transferable to different visual interactive tasks. The evidence comes from comparisons against same-model minimal-harness baselines on three additional benchmarks; the abstract describes the margin as substantial but does not give specific numbers.

Perspective

This work targets researchers and practitioners building multimodal agents that need long-horizon vision and interactive reasoning, in settings where the model can perceive the environment through visual observations and must retain and retrieve past observations. The results described in the abstract center on ARC-AGI-3 and three additional visual game and puzzle benchmarks, so its direct scope is this class of visual interactive environments; the authors position VISTA as a general-purpose visual harness that transfers with minimal adaptation, providing a basis for exploring further visual environments and longer-horizon tasks.

The abstract does not describe VISTA's concrete implementation, the storage and retrieval details of the visual memory, or the names and specific scores of the three additional benchmarks, so the size of those comparisons is known only through the phrase substantially outperforms. The perfect 100.00 and 57.4% action savings on ARC-AGI-3 were obtained on that benchmark against first-time human participants, and transfer to other visual environments still depends on the original experiments. In addition, this summary is based only on the abstract, without the body, figures, or appendix, so questions about ablations, failure cases, and computational cost remain open.

Sources