Controlled distillation study finds on-policy rollouts are not universally better, with KL direction and learning rate dominating
Synopsis
Independently varying rollout policy, token-level KL direction, and learning rate in strong-to-weak distillation across Llama 3 and Qwen2.5 model families and scientific, medical, and arithmetic reasoning tasks, the study finds no consistent advantage for on-policy rollouts in final accuracy, catastrophic forgetting, or parameter-update sparsity; forward KL is robust to rollout policy while reverse KL is substantially more sensitive and favors student-generated rollouts, learning rate governs forgetting and sparsity, and on-policy data helps only for generalisation to harder Countdown variants, an advantage that does not reliably persist after subsequent RLVR.
Interpretation
In a controlled strong-to-weak distillation testbed, independently varying rollout policy, token-level KL direction, and learning rate leaves on-policy and off-policy distillation nearly indistinguishable in final task accuracy, with best mean accuracies of 72% for OnPD and 73% for OffPD. Prior evidence for on-policy advantages came largely from SFT-versus-RLVR comparisons that changed the objective, supervision source and density, and optimisation procedure together; this work decouples rollout policy from KL direction, breaking the conventional pairing of forward KL with off-policy and reverse KL with on-policy learning. Llama-3.1-8B teachers distilled into Llama-3.2-1B students on three datasets (MedReason, Science, Countdown), three random seeds per configuration, 150 full-parameter fine-tuning steps, replicated with Qwen2.5-7B teachers and Qwen2.5-1.5B students.
KL direction shapes task performance and output coverage more clearly than rollout policy: forward KL reaches 71–73% across all rollout policies and learning rates, whereas reverse KL ranges from 35% to 72% and is highly learning-rate sensitive; forward KL also yields larger pass@10 gains under repeated sampling. The work manipulates KL direction as an independent design axis and offers a gradient-level account: the forward-KL derivative with respect to student logits has bounded coordinates and its update difference is bounded linearly by the total-variation distance between rollout distributions, while the reverse-KL derivative is weighted by the student–teacher probability ratio and can grow arbitrarily large. Derivations of token-level KL gradients (Appendix B) and rollout-stability theorems (Appendix C), validated on a continuous student–teacher rollout spectrum where forward KL stays above 80% accuracy and varies by only 5.2 percentage points, while reverse KL fluctuates widely and occasionally collapses at the higher learning rate.
Catastrophic forgetting and parameter-update sparsity are governed mainly by learning rate rather than rollout policy: at the lower learning rate mean OOD performance changes by at most 1.3 percentage points with 85.3–89.6% sparsity, while at the higher rate OOD drops by 11.2–14.0 points and sparsity falls to 51.9–60.0%. This directly tests the existing hypotheses that on-policy learning reduces forgetting and produces sparser updates, and in matched comparisons OffPD produces updates at least as sparse as OnPD. Seven OOD benchmarks (MMLU-Pro, TruthfulQA, IFEval, HumanEval, EQ-Bench, BBQ, ToxiGen) totalling 3,935 examples, plus a learning-rate sweep showing sparsity decreasing approximately linearly with learning rate.
On-policy data does help generalisation to harder tasks: on Countdown-4E, more on-policy rollouts consistently yield higher pass@k under both KL directions, but this initial advantage does not reliably persist after subsequent RLVR, where the strongest sustained performance comes from off-policy checkpoints. The work separates final in-distribution accuracy, output coverage, generalisation to harder variants, and quality as a starting point for later RLVR, showing that the value of on-policy data depends on the evaluation setting. Countdown-3 checkpoints trained with 300 RLVR steps on Countdown-4, three seeds per distillation configuration for 24 runs; additionally, on the longer-rollout Numina–MATH task with 622-token average teacher responses, OnPD with reverse KL achieves the highest MATH-500 accuracy at the lower learning rate, though that experiment has one run per condition and the authors treat it as suggestive.
Perspective
The study targets the controlled strong-to-weak distillation setting, applying to student models of at most 1.5B parameters and reasoning traces of at most about 2,000 tokens, across the Llama 3 and Qwen2.5 families and scientific, medical, and arithmetic reasoning tasks. It lets practitioners judge, given a KL direction and learning rate, whether the extra generation cost of on-policy rollouts is warranted, and it suggests including off-policy distillation as a standard baseline; for suppressing incidental teacher-style transfer, on-policy with reverse KL shows a tendency to preserve the student's English response style.
Readers should still watch several things: the conclusions rest on student models of at most 1.5B parameters and relatively short reasoning traces, so whether larger models and longer rollouts show the same picture remains open; the long-rollout Numina–MATH experiment has one run per condition and the authors themselves treat it as suggestive; the teacher model is held fixed, leaving open whether OnPD–OffPD differences emerge under other student–teacher combinations; and the loaded text is the full body with appendices, but figures appear as references, so specific curve values can only be understood through the prose descriptions in the main text and appendices.
