Skip to main content
Back to timeline
arXivSource publication:

In multi-teacher on-policy distillation, loss averaging and Adam momentum shape capability integration, while top-64 gradient fidelity does not guarantee task gains

Related research and updates

Synopsis

Using Qwen3-1.7B with four RL teachers from the same initialization (mathematics, code, instruction following, science), this study compares gradients, optimizer updates, and task learning curves, with SmolLM3-3B diagnostics: loss averaging implicitly weights longer responses more, Adam's first moment aligns updates across teachers (cosine 0.83 between teachers, 0.96 between averaging rules), BF16 rounding hides about 97% of FP32 master-weight changes and shows only 7-11% of BF16 weights changing, and the top-64 intersection KL gradient closely matches the full-vocabulary gradient yet its mathematics accuracy is 2.6 points above sampled-token policy gradient under response averaging and 2.1 points below under global token averaging.

AI-generated editorial illustration: From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

Interpretation

The averaging rule is itself an implicit weighting over responses and domains: global token averaging assigns domain weight by token share, domain token averaging fixes domain weights but still weights responses by length, and only domain response averaging makes responses equally weighted. The authors derive a covariance identity writing domain token averaging's within-domain gradient as the unweighted domain gradient plus the covariance between length and per-response gradients, showing that Open-MOPD's token-share correction recovers domain token averaging while retaining within-domain length weighting. Domain token and domain response averaging are compared on identical responses at fixed student parameters, with a mean gradient cosine of 0.68 across six batches; the covariance identity is verified in FP64 for all 24 batch-domain groups with a very small maximum relative residual.

Adam's first moment makes parameter updates from different teacher losses point in similar directions even when current gradients are nearly orthogonal, while BF16 rounding makes widespread FP32 parameter movement appear sparse. The authors separate current supervision from optimizer history: update cosine between teachers exceeds 0.83, but resetting the first moment or using norm-matched SGD reduces it to near zero; on domain-specific inputs the update cosine after subtracting the zero-gradient baseline is below 0.01, rising to about 0.24-0.35 with shared inputs. One-step updates are applied to saved Adam states; under domain response averaging about 97% of FP32 master weights differ from initialization versus 7-9% after BF16 rounding (7-11% including domain token and global token averaging), and about one third of parameters account for 90% of squared change in FP32 versus about 4% in BF16.

Vocabulary truncation that more closely matches the full-vocabulary gradient does not consistently improve task performance, and its effect depends on the averaging rule. The top-64 intersection KL retains more than 99.9% of student probability in Qwen and nearly matches the full-vocabulary KL gradient direction, yet relative to sampled-token policy gradient its mathematics accuracy is 2.6 points higher under response averaging and 2.1 points lower under global token averaging, showing gradient fidelity does not predict task-level gains and losses. PG, top-64 intersection, teacher-top-64, and full-vocabulary KL are compared at matched prefixes and Adam states; with 1, 16, or 64 samples per prefix the PG gradient cosine to full KL rises from 0.54 to 0.87 and 0.97 while top-64 exceeds 0.999; on SmolLM3-3B top-64 retains 97.42% of student probability initially and 98.56% after training, yet 8.42% and 3.65% of weighted prefixes retain less than 90%.

With the PG loss, momentum-free SGD achieves higher four-task mean scores than Adam under all three averaging rules. The authors connect this to the possibility that momentum retains directions favored by earlier responses and delays adaptation to current supervision, and note that whether smaller Adam, gradient clipping, or second-moment normalization explain why losses with similar expected gradients produce different task outcomes remains to be tested. Over 500 training steps with the PG loss, domain response averaging gives 39.94 versus 38.91, domain token averaging 39.46 versus 38.63, and global token averaging 39.26 versus 38.79; the mean-score range among the three averaging rules is 0.28 percentage points under Adam and 0.06 under top-64 intersection KL.

Perspective

The work targets post-training practitioners using same-initialization RL teachers for multi-teacher on-policy distillation, at the Qwen3-1.7B scale with four domain teachers (mathematics, code, instruction following, science) and supplementary SmolLM3-3B diagnostics. It provides controlled measurements at fixed student parameters, fixed responses, and saved optimizer states, plus 500-step task curves, so it can directly inform joint choices of response weighting rule, vocabulary candidate budget, and optimizer (Adam versus momentum-free SGD), and it suggests reporting both FP32 master-weight and BF16-weight statistics when assessing parameter change.

Conclusions rest on limited model families and scales, and the authors note observations may differ at larger scale; whether smaller Adam, gradient clipping, or second-moment normalization explain why losses with similar expected gradients yield different task outcomes remains open. SmolLM3 diagnostics show that high mean coverage coexists with low-coverage prefixes and relative gradient errors up to 24.54%, so how such long-tail prefixes affect training under larger vocabularies or more diffuse predictions deserves continued observation. This summary is based on the paper's main text and appendices without the original figures, so some graphical details of individual numbers cannot be checked here.

Sources