Skip to main content
Back to timeline
arXivSource publication:

ReLaViS lifts ScreenSpot-Pro grounding accuracy by 3.1 points to 56.3% via single-round recurrent latent visual search

Related research and updates

Synopsis

The work proposes ReLaViS, which formulates GUI grounding as recurrent latent visual search within a single model interaction: at each step a spatial search head uses the hidden state to query the screenshot's visual tokens, producing a spatial search distribution that aggregates those tokens into latent visual evidence fed back as the next input embedding, supervised by GUI-aware coarse-to-fine trajectories built from flat element annotations; built on Qwen2.5-VL-7B, it outperforms matched single-step baselines on all five benchmarks, raising ScreenSpot-Pro accuracy from 53.2% to 56.3% with only 3.5% more inference FLOPs.

Source-provided article image: Recurrent Latent Visual Search for GUI Grounding
Figure 1 ·

Figure 1: Conceptual comparison of grounding paradigms. ReLaViS performs GUI-aware coarse-to-fine visual search by updating explicit spatial search states within a single model interaction.

arXiv

Interpretation

GUI grounding is reformulated as recurrent latent visual search inside the model, where the search state is an explicit spatial distribution aligned with the visual token grid and multi-step search completes within a single interaction without external tools or image re-encoding. Prior multi-step approaches either rely on textual chains of thought misaligned with visual space or on external visual tools such as zoom-in and cropping that incur multiple interaction rounds and re-encoding; existing latent visual reasoning also leaves the search focus over the visual token grid implicit. The method specifies the spatial search head, latent visual evidence aggregation, and deterministic coordinate readout; across five benchmarks it is compared with matched single-step baselines, reaching 56.3% versus 53.2% on ScreenSpot-Pro with only 3.5% more inference FLOPs.

Coarse-to-fine search trajectories are constructed from flat GUI element annotations, instantiating a GUI-aware inductive bias without scarce ground-truth UI trees. Hierarchy-rich GUI corpora mainly cover mobile or web platforms, and desktop grounding data do not consistently pair UI trees with aligned instructions and target elements; this work recursively partitions candidate element sets and selects key steps by candidate-count reduction to approximate the GUI hierarchy from flat annotations. Trajectory statistics over 105,000 training and validation samples show the mean candidate-set size contracting from 74.25 to 28.01, 10.11, and 1.00 at the default four steps, with 89.68% of samples providing exactly two intermediate key steps; on 1,125 macOS screenshots with UI-tree annotations, set IoU exceeds geometric zoom by 12.7 and 16.8 percentage points.

The gains from recurrent search appear mainly on small targets, and later search steps functionally depend on preceding spatial search states. Single-step grounding struggles with high-resolution screens, small targets, and dense layouts; this work links the gains to target size and tests the causal role of preceding search distributions through interventions. After binning by relative target area, performance is comparable to single-step baselines in the largest-target bins while gains are generally larger in smaller-target bins, with a similar tendency in within-benchmark analyses across five benchmarks; injecting a target-aligned Gaussian at steps 1:3 raises accuracy by 2.0–13.1 percentage points, whereas a target-opposite Gaussian lowers it by 4.7–14.2 percentage points, with the strongest sensitivity at the third step.

Three-stage training and candidate-wise coverage supervision are key design choices for recurrent search, and the gains persist at a smaller model scale. Simply adding recurrent steps or repeating target-element supervision at every step does not yield comparable improvement, indicating that the benefit comes from GUI-aware coarse-to-fine supervision and training–inference adaptation rather than extra computation alone. Four-step variants with final-step-only or every-step target supervision reach 54.0% and 53.9% versus 56.3% for full ReLaViS; removing Stage 3 drops accuracy to 53.0%, removing Stage 2 gives 54.7%, and omitting Stage 1 gives 54.9%; on Qwen2.5-VL-3B average accuracy rises from 58.1% to 61.2%.

Perspective

The result targets GUI grounding settings that take a screenshot and an instruction as input and output a click coordinate, covering desktop, web, and mobile interfaces, and is validated at both Qwen2.5-VL-7B and 3B scales. It enables multi-step visual search within a single model interaction, requiring no element annotations, UI tree, zoom-in or cropping tool, or additional interaction rounds, making it suitable for GUI agent deployments sensitive to latency and inference cost; combined with external zoom-in, the five-benchmark average can rise further to 70.0%.

More search steps are not monotonically better, with four steps optimal under the current setup, and the behavior at larger step counts remains an open question; gains vary markedly across benchmarks, only 0.3 percentage points on ScreenSpot-v2, and are associated with target size. Failure cases show the final click can still drift under limited visual token grid resolution, instruction–annotation ambiguity, and limited ability to distinguish abstract visual semantics. The current work focuses on click actions, and extension to other actions such as drag-and-drop and to general visual search tasks is not yet verified. In addition, equations, tables, and figures appear as placeholders in the body text, so specific numerical details require consulting the original appendices.

Sources