ARR lets a multimodal retriever iteratively rewrite its query from retrieved items, topping averages on M-BEIR and seven zero-shot benchmarks
Synopsis
The authors introduce the AutoRegressive Retriever (ARR), a multimodal retrieval model that alternates between retrieving an item and updating the query embedding, appending each selected item's multimodal content to the query history before re-encoding and using the final embedding to rank the collection; training has two stages, supervised fine-tuning with stepwise contrastive supervision to teach feedback interpretation, then reinforcement learning that treats feedback items as actions and optimizes their selection using the final reciprocal rank of a relevant item, with a query-side adapter enabling this against a fixed item index.
Figure 1: Autoregressive retrieval with ARR. Retrieved items provide context for successive query embeddings; the final embedding ranks the collection.
arXivInterpretation
ARR turns retrieval itself into an autoregressive sequence: at each step it selects an item by embedding similarity, appends that item's content and metadata to the query context, and re-encodes to produce the next query embedding, with the final embedding ranking the collection; item embeddings remain independently encoded, preserving the precomputed-index search interface. Standard universal multimodal retrieval encodes a query once and leaves it fixed, and prior retrieval-feedback methods route retrieved items through auxiliary rewriting or reranking mechanisms; ARR has the retriever itself consume and select feedback items, and its history mask applies only to feedback selection, so previously selected items stay eligible for final retrieval. The paper formalizes states and the selection policy (Eqs. 3-5) and reports results on M-BEIR and seven zero-shot benchmarks; an ablation shows a model not trained to use feedback drops from 55.23 to 46.90 on average when inference is extended from one to four states, indicating feedback use must be learned.
Two training stages separate interpreting feedback from selecting it: SFT applies stepwise contrastive supervision on offline trajectories plus a degradation penalty on relative increases in contrastive loss between consecutive states, while RL freezes the SFT backbone, trains only a query-side LoRA adapter, optimizes item selection with a GRPO-style objective using the final reciprocal rank of a relevant item as reward, and adds contrastive supervision at the final state. SFT contrastive losses supervise each state without assigning credit to actions; ARR's RL action space is retrieved item identities rather than generated text tokens, and the reward comes directly from the final ranking, so the model can favor intermediate items that are not annotated positives but help later retrieval. Ablations show removing the degradation penalty leaves initial performance nearly unchanged (55.69 vs. 55.63) but eliminates the feedback gain (55.68 vs. 56.01 final); removing GRPO yields 55.29 at the final step, below both the full RL model and its SFT initialization; removing the final embedding contrastive loss gives 56.43, still above SFT's 56.01.
On in-domain M-BEIR evaluation, ARR-RL achieves the highest average among compared methods at both scales, 56.7 for 2B and 61.0 for 8B; at 8B it exceeds the strongest baseline average, TRACE's 58.8, by 2.2 points, and RL improves over ARR-SFT by 0.7 and 1.2 points respectively. Gains concentrate on configurations requiring connection across different information structures: for the 8B model, RL raises EDIS text-to-multimodal recall from 66.8 to 73.2 and OVEN multimodal-to-text recall from 58.7 to 63.5; the paper also notes slight degradation on a few configurations including VisualNews at 2B and CIRR at both scales. Results come from the M-BEIR local-pool setting over 16 dataset-task combinations and are compared against MM-Embed, LamRA-Ret, TRACE, and ELVA; the RL stage samples 2K queries from each of the 16 combinations, 32K queries total, with a pool formed from the SFT model's top-100 plus the annotated positive.
In zero-shot evaluation, 8B ARR-RL obtains the highest reported average, 79.01, versus 74.10 for the strongest baseline average, and scores 77.9 on Visual Dialog and 66.0 on Multi-round FashionIQ, indicating transfer to dialogue and interactive retrieval; models from both training stages benefit from feedback at inference time. The paper separates two benefits of feedback: using feedback at inference improves subsequent retrieval (ARR-SFT from 77.95 to 78.39, ARR-RL from 78.42 to 79.01), while learning from feedback during training also improves the initial query embedding, with ARR-SFTN=4 outperforming ARR-SFTN=1 by 0.40 points before any feedback is observed. Zero-shot evaluation covers seven benchmarks, ShareGPT4V, Urban-1K, Flickr30K, CIRCO, GeneCIS, Visual Dialog, and Multi-round FashionIQ, without task-specific fine-tuning; the paper notes several benchmarks reuse COCO or FashionIQ images, so the protocol does not imply all visual content is unseen.
Perspective
The work targets universal multimodal retrieval, where queries and items may be text, images, or combinations, with evaluation spanning M-BEIR's news, fashion, Wikipedia, and miscellaneous domains plus seven zero-shot benchmarks. The method is designed to work against a fixed item index, and RL updates only a query-side LoRA adapter, so it suits systems that already have precomputed item embeddings and want better query understanding without rebuilding the index. The paper reports Iter-2 as a favorable trade-off between test-time scaling and retrieval performance, and uses four feedback steps by default in training and inference; for multi-turn or conversational settings, the Visual Dialog and Multi-round FashionIQ results offer reference points.
The paper reports slight degradation on a few task configurations (VisualNews at 2B, CIRR at both scales) and hypothesizes this may stem from RL training intensively on a relatively small data subset introducing task-specific biases, an explanation that remains to be verified. Feedback gains flatten after Iter-2, and four-state training beats two-state training at Iter-4 by only 0.10 points, so the return on additional inference steps needs observation across more tasks. In the zero-shot protocol several benchmarks reuse COCO or FashionIQ images, so whether all visual content is unseen is not certain. The degradation penalty is described as encouraging but not guaranteeing improved rankings after feedback, leaving its robustness across data distributions an open question.
