Forward KL is proved to inflate student entropy above the teacher, making divergence an implicit entropy regularizer in distillation
Related research and updatesSynopsis
Taking an entropic perspective on knowledge distillation, the work proves that forward KL inflates the student's entropy above the teacher's, with cross-entropy training as a special case yielding an identity verified quantitatively in pretraining and supervised finetuning; reverse KL deflates entropy until the student-teacher gap grows too large, interpolating between the two changes entropy smoothly early in training but abruptly at convergence; the lower entropy of on-policy distillation comes from token-level reverse KL rather than on-policy sampling, so the divergence acts as an implicit entropy regularizer, most clearly in self-distillation where the best divergence hyperparameters compensate for the entropy deflation caused by conditioning on privileged information.
Figure 1: Decoupling divergence from sampling in distillation. A prompt x x is completed by the sampling distribution p p , and the student π θ \pi_{\theta} is then compared to the teacher π ∗ \pi^{*} on each token of the completion through a divergence ℓ \ell between their next-token distributions. Writing the objective at the token level keeps the two choices independent, whereas a sequence-level divergence ties them together. Details in Section 2 .
arXivInterpretation
The paper proves that forward KL inflates the student's entropy above that of the teacher, and notes that cross-entropy training is a special case, yielding an identity about entropy. Distillation properties were not yet well understood; this work links the choice of divergence in the distillation objective directly to the student's entropy level and provides a quantitatively verifiable relation rather than only empirical observation. The evidence is a formal proof, together with quantitative verification of the identity in pretraining and supervised finetuning.
Other divergences carry no such guarantee: reverse KL deflates entropy until the gap between student and teacher gets too large, and interpolating between the two changes entropy smoothly early in training but abruptly at convergence. This complements the forward-KL result on the other side, showing that entropy behavior is not shared by all divergences but depends on the chosen divergence, and characterizing how entropy evolves along the interpolation path across training. The evidence is theoretical analysis of the divergences, together with a description of entropy behavior in early training versus at convergence.
The lower entropy of on-policy distillation comes from token-level reverse KL, not from on-policy sampling. This corrects the intuitive attribution of on-policy distillation's low entropy, locating the cause in the token-level reverse KL term of the objective. The evidence is a decompositional analysis of the entropy source in on-policy distillation, separating the sampling scheme from the divergence form.
The divergence therefore acts as an implicit entropy regularizer, whose role is clearest in self-distillation: as conditioning on privileged information deflates entropy, the divergence hyperparameters that work best are those that compensate for it. The work elevates divergence from a mere objective choice to a mechanism for tuning entropy, and explains the mutual compensation between divergence hyperparameters and privileged-information conditioning in self-distillation. The evidence comes from analysis of the relation between divergence hyperparameters and the entropy deflation caused by conditioning on privileged information in the self-distillation setting.
Perspective
The results are aimed at researchers and engineers who train large language models with distillation, in settings where the distillation objective is defined by a divergence and the student's entropy level matters, including pretraining, supervised finetuning, on-policy distillation, and self-distillation. It offers a set of criteria for how divergence affects entropy: forward KL inflates entropy, reverse KL deflates it, interpolation is smooth early and abrupt at convergence, and in self-distillation divergence hyperparameters can compensate for the entropy deflation caused by conditioning on privileged information. On this basis, readers can treat divergence as an implicit entropy regularizer when selecting distillation objectives and tuning, rather than only as a distance measure for fitting the target.
The text presents only abstract-level information: it does not give the concrete form of the proof, the expression of the identity, the scale and setup of the pretraining and supervised finetuning experiments, the specific conditions under which the interpolation becomes abrupt, or the specific hyperparameter values in self-distillation. Readers who need to judge how far these conclusions apply to their own model scale and data distribution would still need the derivations and experimental details in the full paper.
