LSPD recasts on-policy distillation as KL-regularized RL: +1.59 Avg@16 across six math reasoning benchmarks, with a replay variant matching vanilla OPD in about a quarter of rollout batches
Synopsis
The work reinterprets the reverse-KL objective of on-policy distillation (OPD) as KL-regularized policy optimization and introduces Least-Square Policy Distillation (LSPD), which replaces policy-gradient updates with robust quadratic log-probability matching plus entropy regularization, enabling multiple updates per rollout and reuse of historical trajectories; across six mathematical reasoning benchmarks and three teacher-student settings it improves Avg@16 by +1.59 and Pass@16 by +1.87 on average, and its fully off-policy replay variant LSPD-RB reaches performance comparable to vanilla OPD using roughly one quarter of the rollout batches.
Interpretation
The paper gives an exact rewriting of the OPD reverse-KL objective: with a fixed reference policy, OPD is equivalent to KL-regularized policy optimization whose reward is the teacher-to-reference log-likelihood ratio, not an added regularizer. Prior work had noted a connection between OPD and KL-regularized RL, but this paper states it as a token-level exact equivalence and uses it to import optimistic estimation and off-policy data reuse from value-based RL into distillation. A derivation-level equivalence argument; the text explicitly calls it an "exact reformulation of OPD rather than an additional regularization term" and gives a per-prefix KL chain-rule decomposition.
The resulting LSPD uses a robust (Huber-type) quadratic log-probability matching loss plus an explicit entropy bonus as its practical objective, which naturally supports multiple updates per rollout batch and updates sampled from a replay buffer. Common OPD implementations rely on PPO/GRPO-style policy gradients and are largely on-policy, with limited historical reuse and exploration; LSPD writes the update as matching over accumulated prefix-token pairs, making off-policy optimization a property of the objective itself. The method section specifies the objective, the Huber threshold design, and the algorithm; the ablation shows that removing entropy regularization changes Avg@16 only modestly while Pass@16 drops in every method-benchmark comparison, indicating the entropy term mainly affects solution coverage.
On six mathematical reasoning benchmarks and three Qwen3 teacher-student settings (8B to 4B, 4B to 1.7B, 1.7B to 0.6B), LSPD averages 31.60 Avg@16 and 53.30 Pass@16, beating the strongest baseline EOPD (+0.91/+0.85) and standard OPD (+1.99/+2.27); LSPD-RB raises Pass@16 further to 54.86. Relative to existing distillation baselines, the results report both average generation accuracy and multi-sample coverage, and show stronger scaling as the sampling budget grows (average Pass@32/64 of 58.51/61.94 versus EOPD's 57.00/60.43). Tabulated results over 18 model-benchmark combinations, with LSPD best or tied-best Avg@16 in 11 of 18; the two variants together rank top-2 in all 36 metric-setting combinations; Pass@ scaling is reported up to k=64 on AMC23, AIME24, and AIME25.
The fully off-policy replay variant LSPD-RB saturates within about 10 training steps while baselines need more than 40 steps, and the paper reports it matches vanilla OPD using roughly one quarter of the rollout batches. This turns the off-policy data-reuse advantage of value-based RL into a measurable rollout-budget saving, and shows that more updates per rollout batch improves sample efficiency with diminishing returns beyond a certain number. Training curves and step comparisons come from AMC23, AIME24, and AIME25 with a Qwen3-4B teacher and 1.7B-Base student, collecting 64 prompts with 4 responses each per step; on the theory side, the idealized optimistic formulation is shown to attain a logarithmic reverse-KL regret bound under online exploration.
Perspective
The results target LLM distillation with mathematical reasoning as the training domain, in teacher-student settings where token-level teacher feedback is available and the student and teacher have sufficient coverage (the condition set in Assumption 6.1); experiments use three Qwen3 scale pairings and DAPO-Math-17K training data, with evaluation on MATH-500, Minerva, Olympiad-Bench, AMC23, AIME24, and AIME25, plus out-of-domain checks on GPQA, MMLU, WinoGrande, and HumanEval. For practitioners aiming to cut rollout budgets, the LSPD-RB replay-buffer path offers a reusable engineering option; for theory readers, the logarithmic reverse-KL regret bound of the idealized optimistic formulation provides a contrast with SFT/forward-KL views.
The regret bound applies to the idealized optimistic formulation, whereas the practical objective substitutes a Huber-type robust penalty and an entropy bonus for the pointwise optimistic term, and the text does not give a quantitative bridge between the two; the benefit of off-policy reuse diminishes beyond a certain number of updates, and whether that turning point is consistent across model scales and data distributions remains to be seen; out-of-domain evaluation shows no obvious degradation on GPQA, MMLU, WinoGrande, and HumanEval, but training stays within math data, so preservation of other capabilities needs more evidence; additionally, this evidence bundle is a full-text parse in which some table values appear as placeholders, so specific hyperparameters and individual point values should be checked against the original tables.
