Residual Advantage lets the teacher only redistribute step credit: all 24 comparisons improve, and Co-RA adds 1.0–1.5 points
Synopsis
The authors propose Residual Advantage (RA), which treats the teacher–student probability residual as a bounded one-step reward, subtracts the corresponding state value under the student policy to form an advantage, and centers it within each response before adding it to the verifier advantage, so the verifier sets a response's mean label while the teacher only redistributes credit among its steps; with Qwen3-1.7B/4B-Base students and a Qwen3-8B teacher, RA combined with GRPO or REINFORCE++ improves all 24 three-seed comparisons on three mathematical benchmarks, raising macro Avg@8 by 1.7–3.6 points and Pass@8 by 3.9–6.3 points, surpassing teacher-only OPD; Co-RA updates a teacher LoRA with verifier advantages on the same scored student batch and adds a further 1.0–1.5 Avg@8 points.
Figure 1: Local teacher guidance redistributes the verifier’s response-level credit. (a) At one prefix, the reward Q d ( s , v ) = q s ( v ) − p s ( v ) Q_{d}(s,v)=q_{s}(v)-p_{s}(v) over the student’s likeliest tokens, its value V d p ( s ) V_{d}^{p}(s) under the student policy, and the advantage A s d ( a ) A_{s}^{d}(a) of the sampled token, here the student’s third choice and the teacher’s first. (b) Labels along one response, with blue and coral stems for positive and negative guidance terms δ t \delta_{t} ; each label is A V A^{V} plus a centered guidance term, so the labels’ weighted mean is A V A^{V} .
arXivInterpretation
RA treats the teacher–student probability residual as a bounded one-step reward, subtracts the state value of that reward under the student policy (the student value baseline) to form a local advantage, and centers it within each response before adding it to the verifier advantage, so the guidance term has zero mean within every response and the verifier advantage remains the response's mean label. Unlike OPD, which uses a log ratio or one divergence coordinate, RA judges the sampled token against the other candidates the student would consider, weighted by the student's probability, and explicitly removes the teacher's response mean; the authors prove the local advantage is the first-order decrease of the squared distance along a move toward the teacher, that its expected update is the exact logit gradient of that distance, and that it is bounded by the student–teacher distance times the probability the student leaves to other tokens. Proposition 1 states the three properties with proofs in Appendix A.1; Equation 6 holds for every response and Equation 8 gives per-term and aggregate guidance bounds; Appendix C.1 records the mean absolute guidance declining over a training run at a fixed weight.
With Qwen3-1.7B-Base and Qwen3-4B-Base students, a Qwen3-8B teacher, and two epochs on DAPO-Math-17K, adding RA to GRPO or REINFORCE++ improves all 24 three-seed benchmark means, raising macro Avg@8 by 1.7–3.6 points and Pass@8 by 3.9–6.3 points, and both combinations surpass teacher-only OPD. Prior work largely treats OPD and RLVR as separate paradigms or simple mixtures; here the teacher enters as a critic-like signal inside RL, and at a matched budget it exceeds both OPD followed by REINFORCE++ and a joint reverse-KL combination. Table 1 reports Avg@8 and Pass@8 for two students × two estimators × three benchmarks, with three-seed means and sample SD in Appendix D; weights were selected on a validation split of training prompts, and the evaluation benchmarks played no part in selection.
Co-RA updates a teacher LoRA with the base estimator's verifier advantages on the same scored student batch each iteration and uses the updated teacher in the next iteration's residual, adding 1.0–1.5 macro Avg@8 points to the four RA runs, with Avg@8 rising in all twelve benchmark means. Unlike VPD, which refines a feedback-conditioned teacher with outcome supervision and then distills it, Co-RA applies the base estimator's clipped update directly to a teacher LoRA so the teacher adapts to the reasoning along the student's own solution paths; the adapter starts at zero, so the first student update coincides with RA. Table 1 and Appendix D.9 show each adapted teacher scores below the frozen teacher on its own while the students it guides improve, indicating a better guide rather than a better solver; coverage changes are mixed, with Pass@8 rising in four benchmark means, unchanged in two, and falling in six.
Ablations show the student value baseline and response centering each contribute, and the probability residual outperforms the log ratio: removing the student value baseline, removing centering, using the raw residual, or using a direct squared-distance loss all lower performance, and replacing the residual by the log ratio also trails full RA. The authors place the residual and the log ratio at the two ends of one Bregman family and note that only the quadratic endpoint has bounded teacher influence; offline statistics show the residual puts 59% of its guidance on the 12% of tokens where the student is choosing among live alternatives, whereas the centered log ratio puts 50% of its total on the 4% of tokens the teacher assigns very low probability. Table 2 compares controls at weights 5, 10, and 20; Table 3 compares log-ratio variants; Appendix C.2 computes guidance placement over 256 offline responses and 176,628 tokens; Appendix C.3 compares the sampled label and the direct loss in update share.
Perspective
The results are aimed at researchers and engineers post-training reasoning models with verifiable rewards, in settings where student and teacher share a vocabulary, the teacher supplies next-token distributions in non-thinking mode, and an answer checker exists for mathematical reasoning. The method can be layered directly onto GRPO or REINFORCE++, adding one teacher forward pass per rollout, with Co-RA additionally training a teacher LoRA; within the tested budget the authors report that RA's macro Pass@K margin over the untrained student grows with K up to Pass@512.
A careful reader would still watch: validation covers only mathematical reasoning benchmarks and a same-vocabulary teacher, leaving other domains, heterogeneous teachers, or longer training budgets unaddressed; coverage changes are mixed, with Co-RA lowering Pass@8 on some benchmarks; the adapted teacher is a worse solver than the frozen one, showing that guide quality and solver quality are not the same thing; the authors note the teacher-peak branch alters several factors at once and does not isolate a mechanism, leaving open how a dense, critic-like signal should be shaped; wall-clock and peak-memory measurements are not reported.
