Skip to main content
Back to timeline
arXivSource publication:

S²D-OPD keeps only the top 10% of states per response by teacher–reference JSD, lifting held-out accuracy from 46.36% to 47.31% in seven of eight teacher–student settings

Synopsis

The work shows that Direct-OPD's token-level log-ratio reward measures only relative change and can stay fixed while the probability mass that both checkpoints assign to the student's candidate tokens vanishes, whereas the teacher–reference JSD and both KL directions vanish with that mass; it therefore proposes S²D-OPD, which ranks student-sampled states by teacher–reference JSD and masks Direct-OPD supervision at low-divergence states, retaining only the top 10% per response, improving held-out accuracy over dense Direct-OPD on AIME and HMMT benchmarks in seven of eight settings across two teacher pairs and four student models from 1.7B to 8B parameters and matching it in the eighth, without extra forward passes.

AI-generated editorial illustration: Not Every Token Is Worth Distilling: Selective Supervision for Direct-OPD

Interpretation

Through an exact construction, the paper shows that the Direct-OPD reward and its local gradient on the student can remain fixed while the Jensen–Shannon divergence and both KL directions between the teacher and reference checkpoints vanish with the probability mass on the student's candidates; it also gives a bound showing that a small JSD limits how much the teacher's behavior changed. Direct-OPD and Proxy-OPD treat the token-level log-ratio between post-RL and pre-RL checkpoints as dense supervision, implicitly assuming every state's policy shift is worth distilling; this work makes the insensitivity of the log-ratio to absolute probability mass into a provable construction and notes that low JSD means the teacher barely changed. The evidence is the exact construction in Section 4.1 and Proposition 1, proved in Appendix B, together with a JSD bound derived from Pinsker's inequality; this is an analytical result that does not depend on experimental data.

It proposes S²D-OPD, which scores student-sampled states by teacher–reference JSD, retains only the highest-scoring 10% of states within each response, and removes both the reward term and the student-anchor KL term at masked states; the score reuses the teacher and reference probabilities Direct-OPD already computes, so no extra forward passes are needed. Unlike existing selective-distillation criteria that read the student alone, the teacher alone, or teacher–student disagreement, this criterion reads the divergence between the pre-RL and post-RL teacher checkpoints, that is, the size of the policy shift itself. The method is given in Section 4.2, including the JSD definition over the top-k candidate set plus a residual token, the per-response selection rule, and the masked objective; implementation details are in Appendix A.

Across two teacher pairs (R1-Distill-1.5B to JustRL-1.5B and Nemotron-1.5B to QuestA-1.5B) and four students (Qwen3-1.7B/4B/8B and R1-Distill-7B), retaining 10% of states beats dense Direct-OPD in seven of eight settings and ties in the eighth on three held-out benchmarks (AIME 2026, HMMT November 2025, HMMT February 2026); a paired problem bootstrap and exact one-sided sign-flip test over 93 held-out problems give a mean held-out accuracy gain from 46.36% to 47.31%, 0.95 points (95% CI 0.40–1.54). Evidence was previously lacking on whether every per-state supervision signal in dense Direct-OPD is useful; this work provides held-out comparisons across teacher pairs and student scales and reports that under the JustRL pair S²D-OPD matches or exceeds dense supervision over the first 60 steps and is more stable later. The evidence is the controlled comparison in Table 1 and Figure 2, evaluated with Avg@32, selected on AIME 2024/2025 and tested on three benchmarks released after every student model; the authors state in Appendix F that each configuration is trained once and that the confidence interval resamples held-out problems rather than training seeds.

Controlled analyses show dense supervision is largely redundant: a random 10% of states reaches 51.2 peak validation accuracy versus 51.3 for dense Direct-OPD; at the same 10% budget, training on JSD percentile bins improves broadly with percentile, the top bin reaching 52.9 while the three lowest bins fall below the initial student and the two lowest collapse; across teacher pairs, the lower-divergence QuestA pair benefits more from masking (1.0–1.7 points versus 0.0–1.2 for JustRL). This separates how many states are retained from which states are retained: the gain comes from filtering low-divergence supervision rather than from concentrating supervision on states with higher absolute divergence. The evidence is the binning experiment in Section 5.3, which fixes the teacher pair, student, data, and hyperparameters and varies only which states enter the objective (Figure 3), plus the cross-teacher-pair comparison in Figure 4; Appendix C offers a top-k overlap mechanism and the Appendix E cases are hand-selected to illustrate how the score behaves rather than to estimate frequencies.

Perspective

The result targets weak-to-strong policy transfer in mathematical reasoning with verifiable rewards: teacher-side checkpoints are public so no RL is run by the authors, students range from 1.7B to 8B, teachers are 1.5B, and evaluation uses three held-out benchmarks (AIME 2026, HMMT November 2025, HMMT February 2026) released after every student model. What carries over directly is that scoring reuses the teacher and reference probabilities Direct-OPD already computes, so no extra forward passes are needed, and that retention rates from 5% to 20%, per-response or batch-level selection, and JSD or forward-KL scoring all give similar results, making the selection step easy to swap into comparable distillation pipelines. The authors list adaptive retention rates and domains beyond mathematical reasoning as natural next steps.

JSD measures how much the teacher changed relative to the reference, not whether the change is correct: the Appendix E case shows a shift that raises the probability of an erroneous answer (JSD 0.1197, above that response's 0.0744 retention cutoff) is retained just the same, so the quality of retained supervision still depends on the teacher's RL outcome. The score depends on the student only through its top-k candidate set, so a shift the student already follows consumes the same budget as one it has not learned. On scale, teachers are 1.5B and students up to 8B, whereas weak-to-strong transfer is most valuable for much larger students; all experiments use mathematical reasoning with verifiable rewards, leaving code generation or agentic tasks, where RL-induced shifts may be distributed differently, untested. The authors state in Appendix F that each configuration is trained once and that the confidence interval resamples held-out problems rather than training seeds, so per-setting comparisons would be strengthened by repeated runs. In addition, some formulas and figures in the loaded text appear as placeholders, so details tied to specific numeric values follow the prose.

Sources