Skip to main content
Back to timeline
arXivSource publication:

How Retrieval Credit Is Assigned Decides Search-Agent Training Gains: Local Credit Lifts Qwen3-4B Average F1 from 51.25 to 54.34

Synopsis

The study builds corpus-grounded QA data with intermediate evidence annotations and compares three retrieval signals (evidence coverage Cov, coverage with dependencies Cov-Dep, answer match AM) under scalar reward bonuses versus action-aligned local credit, training Qwen3-4B on seven QA benchmarks and finding local credit outperforms scalar, with Cov-local reaching 54.34 average F1 versus 51.25 for outcome-only training.

Source-provided article image: Beyond Outcome Rewards: Constructing and Assigning Retrieval Credit for Search Agents
Figure 1 ·

Figure 1: Observed coverage and schematic credit assignment. (a) An observed search matches v 1 v_{1} and v 3 v_{3} ; the missing father link at v 2 v_{2} blocks dependency credit for v 3 v_{3} , giving Cov 0.67 and Cov-Dep 0.33. Text is abridged. (b) Scalar changes the shared advantage; local retains the outcome advantage and adds credit on executed query tokens (Equations 4 – 5 ). Tool observations are masked. Colour shows support, not magnitude.

arXiv

Interpretation

The same retrieval signal behaves differently under different credit assignment: local credit outperforms scalar incorporation for every signal, with the largest gain for evidence coverage, where Cov-local reaches 54.34 average F1 versus 51.25 for outcome-only training (OO). Prior work largely folds intermediate retrieval rewards into the outcome reward; this work treats what earns reward and how that credit enters the policy update as two coupled design dimensions compared separately. Qwen3-4B is trained on the same 14,000 questions under matched GRPO settings and budget, and 13 conditions are compared across seven benchmarks totaling 3,197 questions, reporting final-checkpoint F1 and EM.

The effect of dependency constraints reverses with incorporation: removing dependency closure raises local F1 from 52.96 to 54.34, while scalar F1 moves from 52.82 to 52.64, showing a more constrained evidence signal does not necessarily provide more useful supervision. Cov and Cov-Dep share the same evidence nodes and grounding and differ only in whether prerequisite nodes must be credited first, isolating the dependency choice. Both coverage signals use the same passage-matching rule and annotations, compared on the same questions and grounding.

Local credit needs to stay aligned with the search that produced the evidence: permuting event masses lowers average F1 by 2.12 for Cov and 3.12 for AM, with larger gaps on multi-hop QA. The permutation conditions preserve signed event masses and support while breaking event-action correspondence, testing whether alignment itself matters. Aligned and permuted conditions are compared on the same questions at the same search step; Cov-permute still exceeds OO (52.22 vs 51.25) while AM-permute falls below it (50.39).

The value of local credit is not confined to groups where the outcome reward cannot distinguish trajectories: for Cov, excluding tied-outcome groups costs 0.32 while retaining only those groups costs 2.27; AM follows the same ordering with losses of 1.44 and 2.19. The flat-only and no-flat restrictions change which groups receive local credit while retaining the outcome advantage, testing whether its benefit concentrates in outcome ties. Using all groups gives the highest aggregate F1 for both signals, and both restrictions also reduce the total amount of local credit.

Perspective

The framework targets multi-step retrieval training for multi-hop question answering, in settings with verifiable answers and corpus-grounded annotations; the authors state that the analysis focuses on evidence acquisition and retrieval credit, and suggest extending the controls across model sizes and retrievers to test whether these credit-assignment behaviors generalize. For readers seeking to reproduce, the paper states that code, trained models and the constructed dataset will be released upon acceptance, with appendices covering training settings, data construction and grounding checks.

The development analysis shows Cov-local raises mean coverage (27.51% to 31.97%) and reduces zero-coverage questions (616 to 551), yet F1 at the same coverage level is not consistently higher, and the question sets at each coverage level differ, so whether evidence use improves remains an open question. The authors also note that the interaction between reward construction and update strength, and independent variation of event timing and reward magnitude, warrant closer study. In addition, the development split overlaps with training evidence (507 questions share some annotated evidence, 693 have fully unseen evidence), which the authors state limits its value for measuring generalization.

Sources