Skip to main content
Back to timeline
arXivSource publication:

A multi-step query-rewriting agent iterates retrieval with corpus feedback, lifting TopiOCQA MRR from 37.8 to 41.4

Synopsis

The work recasts conversational query rewriting as a sequential retrieval problem in which an agent chooses among intent resolution, lexical reformulation, and pseudo-document synthesis, conditioning each rewrite on passages returned by the previous step, and is trained by supervised fine-tuning plus reinforcement learning against retrieval quality alone with no human rewrite annotations; it achieves the highest MRR and NDCG@3 on TopiOCQA and QReCC and transfers zero-shot to TREC CAsT and to dense retrieval backends.

Source-provided article image: Learning Multi-Step Query Rewriting via Corpus Feedback for Conversational Search
Figure 1 ·

Figure 1: Retrieval quality on TopiOCQA and QReCC (BM25) as the 5-step agent is truncated at inference to a maximum of one to five steps.

arXiv

Interpretation

Single-step rewriting is replaced by a multi-step, corpus-feedback sequential decision process in which the agent selects among intent resolution, lexical reformulation, pseudo-document synthesis, and a stop action, conditioning each step on previously retrieved passages. Prior retrieval-aligned methods such as ConvSearch-R1 fix the rewrite before seeing any retrieved document, and pseudo-relevance feedback methods hand-design the feedback mechanism; here the agent learns from retrieval reward which reformulation to apply at each step and when to stop. Compared on TopiOCQA and QReCC with BM25 against five trained rewriters under shared evaluation splits, retriever, and metrics, with ConvSearch-R1's released inference code and weights re-run and its reported scores reproduced.

The full multi-step, multi-action agent reaches MRR 41.4 and NDCG@3 40.6 on TopiOCQA and MRR 57.2 and NDCG@3 55.8 on QReCC, the highest among compared methods, with gains concentrated at the top of the ranking. The single-step, single-action configuration reproduces ConvSearch-R1-level performance, while single-step, multi-action reduces it, showing the gain comes from combining iteration with typed rewriting rather than from the backbone or the action space alone. Controlled configurations isolate backbone, action space, and iteration; the advantage of iteration is absent after supervised fine-tuning and emerges only once reinforcement learning optimizes against the retriever.

Through retrieval-reward optimization alone, the policy converges to a consistent action order: intent resolution first, optional lexical refinement, then pseudo-documents grounded in passages retrieved by earlier steps. Neither the prompt nor the reward enforces this order; SFT rollouts average 1.7 search steps and only 60% open with intent resolution, whereas under RL those alternative openings disappear and episodes lengthen to 4.9 search steps on TopiOCQA and 2.9 on QReCC. Median 8-gram containment of pseudo-documents is 0, indicating no verbatim copying; unigram Jaccard with context passages excluding stop-words averages 0.137 versus 0.014 for an unrelated query's context, and replacing contextual passages with those from an unrelated query drops TopiOCQA NDCG@3 from 40.6 to 25.3.

The policy transfers zero-shot to TREC CAsT 2019 and 2020 and remains effective when the retrieval backend is changed at inference time. The TopiOCQA-trained model exceeds ConvSearch-R1 on NDCG@3 in both CAsT years, and on TopiOCQA both ANCE and Qwen3-Embedding-0.6B improve over BM25 on every metric. Transfer evaluation involves no additional training; on QReCC neither dense backend surpasses BM25, consistent with prior work reporting BM25 ahead on that benchmark, indicating absolute performance still depends on how well the retriever suits the dataset.

Perspective

The result applies to conversational retrieval settings supervised by passage relevance labels and interfaced through an off-the-shelf retriever: training and main evaluation use TopiOCQA and QReCC, CAsT is used only for zero-shot evaluation, and no retriever is trained or adapted in any experiment. The method suits conversational systems that must turn context-dependent user turns into standalone queries and can accept multi-round retrieval cost; the authors note that adding step penalties or latency costs to the reward could yield cost-aware policies that terminate early when conversational ambiguity is low, and that issuing multiple reformulations in parallel per step and rewarding their fused ranking could broaden recall while shortening the sequential loop.

Several open questions remain for a careful reader: the advantage of iteration appears only after reinforcement learning, so reproducibility depends on the specific SFT distillation and GRPO configuration; pseudo-document grounding is measured by 8-gram containment and unigram Jaccard, leaving open whether trivial re-ranking is fully excluded; the cumulative generation cost of sequential interaction is dataset-dependent, with 93% of TopiOCQA trajectories issuing a fifth rewrite versus 4% on QReCC, and the cost-benefit tradeoff is not yet part of the reward; and on QReCC dense retrievers do not surpass BM25, indicating transfer depends on retriever-dataset fit. In addition, equations and some symbols are missing from the parsed text, so training details should be checked against the original appendix for reproduction.

Sources