Public articles linked to the same research event.
arXiv The work proposes ReLaViS, which formulates GUI grounding as recurrent latent visual search within a single model interaction: at each step a spatial search head uses the hidden state to query the screenshot's visual tokens, producing a spatial search distribution that aggregates those tokens into latent visual evidence fed back as the next input embedding, supervised by GUI-aware coarse-to-fine trajectories built from flat element annotations; built on Qwen2.5-VL-7B, it outperforms matched single-step baselines on all five benchmarks, raising ScreenSpot-Pro accuracy from 53.2% to 56.3% with only 3.5% more inference FLOPs.
The work proposes ReLaViS, which formulates GUI grounding as recurrent latent visual search within a single model interaction: at each step a spatial search head uses the hidden state to query the screenshot's visual tokens, producing a spatial search distribution that aggregates those tokens into latent visual evidence fed back as the next input embedding, supervised by GUI-aware coarse-to-fine trajectories built from flat element annotations; built on Qwen2.5-VL-7B, it outperforms matched single-step baselines on all five benchmarks, raising ScreenSpot-Pro accuracy from 53.2% to 56.3% with only 3.5% more inference FLOPs.
The work proposes ReLaViS, which formulates GUI grounding as recurrent latent visual search within a single model interaction: at each step a spatial search head uses the hidden state to query the screenshot's visual tokens, producing a spatial search distribution that aggregates those tokens into latent visual evidence fed back as the next input embedding, supervised by GUI-aware coarse-to-fine trajectories built from flat element annotations; built on Qwen2.5-VL-7B, it outperforms matched single-step baselines on all five benchmarks, raising ScreenSpot-Pro accuracy from 53.2% to 56.3% with only 3.5% more inference FLOPs.
The work proposes ReLaViS, which formulates GUI grounding as recurrent latent visual search within a single model interaction: at each step a spatial search head uses the hidden state to query the screenshot's visual tokens, producing a spatial search distribution that aggregates those tokens into latent visual evidence fed back as the next input embedding, supervised by GUI-aware coarse-to-fine trajectories built from flat element annotations; built on Qwen2.5-VL-7B, it outperforms matched single-step baselines on all five benchmarks, raising ScreenSpot-Pro accuracy from 53.2% to 56.3% with only 3.5% more inference FLOPs.