Public articles linked to the same research event.
arXiv Taking an entropic perspective on knowledge distillation, the work proves that forward KL inflates the student's entropy above the teacher's, with cross-entropy training as a special case yielding an identity verified quantitatively in pretraining and supervised finetuning; reverse KL deflates entropy until the student-teacher gap grows too large, interpolating between the two changes entropy smoothly early in training but abruptly at convergence; the lower entropy of on-policy distillation comes from token-level reverse KL rather than on-policy sampling, so the divergence acts as an implicit entropy regularizer, most clearly in self-distillation where the best divergence hyperparameters compensate for the entropy deflation caused by conditioning on privileged information.
Taking an entropic perspective on knowledge distillation, the work proves that forward KL inflates the student's entropy above the teacher's, with cross-entropy training as a special case yielding an identity verified quantitatively in pretraining and supervised finetuning; reverse KL deflates entropy until the student-teacher gap grows too large, interpolating between the two changes entropy smoothly early in training but abruptly at convergence; the lower entropy of on-policy distillation comes from token-level reverse KL rather than on-policy sampling, so the divergence acts as an implicit entropy regularizer, most clearly in self-distillation where the best divergence hyperparameters compensate for the entropy deflation caused by conditioning on privileged information.
Taking an entropic perspective on knowledge distillation, the work proves that forward KL inflates the student's entropy above the teacher's, with cross-entropy training as a special case yielding an identity verified quantitatively in pretraining and supervised finetuning; reverse KL deflates entropy until the student-teacher gap grows too large, interpolating between the two changes entropy smoothly early in training but abruptly at convergence; the lower entropy of on-policy distillation comes from token-level reverse KL rather than on-policy sampling, so the divergence acts as an implicit entropy regularizer, most clearly in self-distillation where the best divergence hyperparameters compensate for the entropy deflation caused by conditioning on privileged information.
Taking an entropic perspective on knowledge distillation, the work proves that forward KL inflates the student's entropy above the teacher's, with cross-entropy training as a special case yielding an identity verified quantitatively in pretraining and supervised finetuning; reverse KL deflates entropy until the student-teacher gap grows too large, interpolating between the two changes entropy smoothly early in training but abruptly at convergence; the lower entropy of on-policy distillation comes from token-level reverse KL rather than on-policy sampling, so the divergence acts as an implicit entropy regularizer, most clearly in self-distillation where the best divergence hyperparameters compensate for the entropy deflation caused by conditioning on privileged information.