ZJU team runs 25 same-family on-policy distillation pairs on Qwen2.5 from 0.5B to 14B, finding peak capability is predictable from student scale and teacher score, and that weaker teachers can teach stronger students
Synopsis
The study systematically characterizes the scaling properties of same-family on-policy distillation (OPD) on Qwen2.5 (0.5B–14B) math reasoning, finding a regular early useful-transfer regime in which held-out accuracy rises approximately linearly with the square root of token-level reverse KL from the student initialization, and fitting power laws in student scale, teacher scale, and teacher gold score that predict peak accuracy and transfer rate, with every weak-to-strong student peaking above its own teacher.
Interpretation
OPD training dynamics form a two-phase process: a regular useful-transfer regime in which held-out accuracy (gold score) rises approximately linearly in the KL coordinate, followed by a noisy tail of attenuated improvement, saturation, or regression. Prior OPD analyses examined axes such as rollout length and token-overlap ratio, or focused on objective design; this work indexes OPD trajectories by KL and systematically tests their regularity, with linear fits to the first 30 checkpoints of 25 Vanilla-OPD runs obtaining R² around 0.97 and RMSE around 0.30 in accuracy units. 25 teacher–student combinations (weak-to-strong, same-base, strong-to-weak) over Qwen2.5 0.5B–14B, a mixed GSM8K and MATH training split of 14.8K examples and test split of 6.3K, each OPD run lasting at most ten epochs (580 updates) with periodic held-out evaluation; the authors also give a local KL-geometry derivation showing gold score changes to first order while KL changes to second order, so square-root KL linearizes local transfer.
Peak capability and useful-transfer rate follow joint power laws in student scale, capped teacher scale, and measured teacher score, with peak student error nearly proportional to the teacher's remaining error. Unlike reading scaling off parameter counts alone, the joint law adds measured teacher score as a covariate; the two are collinear on the RL-endpoint grid (correlation about 0.99) and cannot be separated there, so the authors use OPD-product teachers in bootstrapped chains whose scores sit below the size trend to break the collinearity and identify the teacher exponents. Under leave-one-scale-out, the joint law cuts RMSE from 3.43 to 1.66 for Vanilla-OPD and from 2.47 to 0.82 for Delta-OPD; extrapolation to the held-out largest student and teacher stays within 0.7 and 0.4 accuracy points; a controlled comparison not entering any fit distills the 7B student from an intermediate 3B checkpoint scoring 66.0, where the joint law predicts a peak of 74.1 against an observed 73.8 and correctly orders it below the 1.5B endpoint teacher's 77.5.
At a matched score, smaller teachers transfer better, so a teacher's own score alone does not define its supervision value. This explains why the 0.5B student peaks at 40.8% with the 3B teacher but only 39.0% and 37.7% under the 7B and 14B teachers, and why the 1.5B OPD product scoring 53.0% teaches the 3B student worse than the 0.5B RL expert scoring 39.8%. The peak law's teacher remaining-error exponent is about 0.97 (Vanilla-OPD) and 1.02 (Delta-OPD), with student-scale exponents of about 0.30 and 0.34 and a negative teacher-scale exponent; five Vanilla-OPD and three Delta-OPD bootstrap-chain cells supply the collinearity-breaking evidence, with all bootstrap intervals excluding zero.
Design choices materially change transfer: Delta-OPD attains a larger matched-KL local gain slope in 15 of 17 shared pairs and a larger peak gain in 12, concentrated in weak-to-strong pairs; pure OPD beats pure off-policy distillation in every cell, an off-policy cold start grows increasingly harmful with the capability gap, and bootstrapping weak-to-strong OPD does not improve on direct transfer from the smallest expert. This reverses bootstrapping gains reported for weak-to-strong fine-tuning in a chess puzzle setting and dissociates a teacher's score from its teaching value at the trajectory level. The cold start costs 6.6, 15.7, and 19.4 points for the 3B, 7B, and 14B students taught by the 0.5B expert, with one epoch of SFT erasing 32 points of the 14B student's initialization that later OPD does not recover; pure OffPD trails pure OPD by 28.4 points in the extreme weak-to-strong cell and by less than one point in same-base cells; bootstrapped chains peak at 60.4 versus 61.4 at 3B, 70.0–71.3 versus 72.7 at 7B, and 76.8 versus 77.7 at 14B.
Perspective
The results are aimed at research and engineering teams doing post-training within a single model family: under Qwen2.5 0.5B–14B, a mixed GSM8K and MATH math-reasoning setting, a 2K-token response cap, and fixed SFT and GRPO recipes, peak accuracy and useful-transfer rate can be estimated in advance from student scale, capped teacher scale, and measured teacher score, while transfer extent behaves as an approximately scale-free KL budget (median 0.30 for Vanilla-OPD and 0.29 for Delta-OPD). This makes a pipeline that trains an RL expert once at small scale and transfers it across scales predictable, and indicates that returns to teacher scaling taper off near the student's own scale.
The evidence is limited to one pretrained family, one math reasoning mixture, one canonical run per teacher–student pair, and aggregate rather than per-prompt evaluation records, so there are no prompt-bootstrap or optimization-seed intervals; the authors also note that parameter count confounds model, data, and compute scale, so the power laws should be read at compute-optimal settings. The rate law does not outperform a raw-affine baseline under leave-one-scale-out RMSE, the Delta-OPD endpoint law rests on only five observed departures, and the Delta-OPD extent exponents are not identifiable, all flagged as exploratory. The length of the linear window, the heterogeneous tails that follow, and the power-law forms of the fitted coefficients remain empirical regularities without a mechanism; recipe constants such as the rollout-length cap, the KL estimator, validation frequency, and the same-step logging convention are held fixed and their influence is untested. Distillation can also propagate a teacher's spurious knowledge, hallucinations, biases, and unsafe patterns, and held-out accuracy provides only partial safeguards.
