Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

VHOP-Router hands multi-step visual retrieval to the embedding model, lifting retrieval from under 5% to 76.3%

The authors introduce VHOP, a data generation framework and benchmark with five core difficulty levels, and use it to train VHOP-Router, an end-to-end pipeline that turns a standard embedding model into an autoregressive multi-step retriever operating in visual latent space and returning linked image chains in a single tool call; experiments report retrieval performance rising from under 5% to 76.3%, agentic search task success up 52.7%, average token length down 61% from 1886 to 728, and versus a strong baseline retrieving the top 50 results per step, 23x fewer in-context images and 35x lower cumulative API payload, with robust generalization to unseen difficulty levels and realistic test sets.