Skip to main content
Back to timeline
arXivSource publication:

Seek's test-time iterative retrieval lifts Qwen2.5-7B to an 82% relative gain over BM25 on BRIGHT and GPT-4.1 to 37.4 nDCG@10

Synopsis

The work introduces Seek, a training-free self-evaluative exploration framework for knowledge retrieval that performs iterative corpus interaction at test time—an LLM generates pseudo-passages conditioned on accumulated relevance feedback, a retriever surfaces fresh candidates, and a dedicated assessor assigns graded relevance judgments that guide subsequent rounds—matching trained rerankers in ranking quality while consistently improving Recall@100 over single-pass BM25 on TREC Deep Learning, and on the reasoning-intensive BRIGHT benchmark achieving an 82% relative gain over BM25 with Qwen2.5-7B, surpassing all trained baselines, and reaching 37.4 average nDCG@10 with GPT-4.1, exceeding the strongest baseline by 37%.

Source-provided article image: Seek: Self-Evaluative Exploration for Knowledge Retrieval
Figure 1 ·

Figure 1. Overview of Seek . Each round generates feedback-conditioned pseudo-passages, retrieves fresh candidates, and assigns graded relevance judgments that inform subsequent generation.

arXiv

Interpretation

Seek is a training-free framework that mitigates the unrecoverability of single-pass retrieval through iterative corpus interaction at test time. Existing LLM-based retrievers and rerankers interact with the corpus in a single pass and commit to the resulting candidate set, so a missed relevant document is permanently unrecoverable; Seek turns that interaction into a multi-round loop. Method description at the abstract level; no round limit, stopping condition, or computational cost details are given.

Each round has an LLM generate pseudo-passages conditioned on accumulated relevance feedback, a retriever surface fresh candidates, and a dedicated assessor assign graded relevance judgments that guide the next round. It closes a generate-retrieve-assess loop and drives exploration with graded judgments rather than a binary signal. Component-level description from the abstract; how the assessor is trained or prompted and the granularity of graded judgments are not expanded in the visible text.

On TREC Deep Learning, Seek matches trained rerankers in ranking quality while consistently improving Recall@100 over single-pass BM25. A training-free method matches rerankers that require training on ranking quality, while recall improvement is reported as a consistent gain. Result statement from the abstract; no specific Recall@100 values, names of compared rerankers, or significance tests are provided.

On the reasoning-intensive BRIGHT benchmark, Seek with Qwen2.5-7B achieves an 82% relative gain over BM25 and surpasses all trained baselines, while Seek with GPT-4.1 reaches 37.4 average nDCG@10, exceeding the strongest baseline by 37%. It extends the benefit of training-free iterative retrieval from conventional passage ranking to reasoning-intensive retrieval, and exceeds trained baselines in that setting. Two concrete numbers from the abstract (82% relative gain, 37.4 nDCG@10 with a 37% margin), but the baseline set, dataset subsets, and variance are not stated.

Perspective

The result targets information retrieval researchers and system practitioners who want to trade additional test-time compute for recall and ranking quality without paying the cost of training a reranker; the abstract positions the framework as training-free and presents evidence on two benchmarks, TREC Deep Learning and BRIGHT, so the scope of the conclusions should be read as those two evaluation settings. For a reader, this means Seek can be treated as a template for test-time iterative exploration, useful for examining whether multi-round interaction is worth replacing the combination of single-pass retrieval plus a trained reranker.

The visible text is only the arXiv abstract page, without body, figures, or experimental detail, so it is not possible to judge which baselines, dataset subsets, and variance underlie the 82% relative gain and the 37.4 nDCG@10; the number of iterations and latency cost, the reliability of the assessor's graded judgments, and behavior of the training-free setting at different corpus scales are open questions a reader must confirm in the original.

Sources