Public articles linked to the same research event.
arXiv The work proposes Segment-wise On-Policy Distillation (Seg-OPD): after controlled reasoning interventions show that replacing student segments with teacher redrafts improves subsequent reasoning accuracy, it selects student segments via an uncertainty metric, obtains corresponding teacher redrafts, and trains the student to prefer teacher redrafts over their paired student segments while retaining dense token-wise on-policy distillation supervision; on mathematical reasoning and competitive programming tasks, Seg-OPD-trained students achieve higher revision success rates than baselines and consistently outperform the compared state-of-the-art baselines in reasoning accuracy, with an average relative improvement of 5.22% across diverse models and tasks.
The work proposes Segment-wise On-Policy Distillation (Seg-OPD): after controlled reasoning interventions show that replacing student segments with teacher redrafts improves subsequent reasoning accuracy, it selects student segments via an uncertainty metric, obtains corresponding teacher redrafts, and trains the student to prefer teacher redrafts over their paired student segments while retaining dense token-wise on-policy distillation supervision; on mathematical reasoning and competitive programming tasks, Seg-OPD-trained students achieve higher revision success rates than baselines and consistently outperform the compared state-of-the-art baselines in reasoning accuracy, with an average relative improvement of 5.22% across diverse models and tasks.
The work proposes Segment-wise On-Policy Distillation (Seg-OPD): after controlled reasoning interventions show that replacing student segments with teacher redrafts improves subsequent reasoning accuracy, it selects student segments via an uncertainty metric, obtains corresponding teacher redrafts, and trains the student to prefer teacher redrafts over their paired student segments while retaining dense token-wise on-policy distillation supervision; on mathematical reasoning and competitive programming tasks, Seg-OPD-trained students achieve higher revision success rates than baselines and consistently outperform the compared state-of-the-art baselines in reasoning accuracy, with an average relative improvement of 5.22% across diverse models and tasks.
The work proposes Segment-wise On-Policy Distillation (Seg-OPD): after controlled reasoning interventions show that replacing student segments with teacher redrafts improves subsequent reasoning accuracy, it selects student segments via an uncertainty metric, obtains corresponding teacher redrafts, and trains the student to prefer teacher redrafts over their paired student segments while retaining dense token-wise on-policy distillation supervision; on mathematical reasoning and competitive programming tasks, Seg-OPD-trained students achieve higher revision success rates than baselines and consistently outperform the compared state-of-the-art baselines in reasoning accuracy, with an average relative improvement of 5.22% across diverse models and tasks.