Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Seg-OPD Trains Students to Revise Reasoning via Teacher Redrafts, Averaging a 5.22% Relative Gain on Math and Competitive Programming

The work proposes Segment-wise On-Policy Distillation (Seg-OPD): after controlled reasoning interventions show that replacing student segments with teacher redrafts improves subsequent reasoning accuracy, it selects student segments via an uncertainty metric, obtains corresponding teacher redrafts, and trains the student to prefer teacher redrafts over their paired student segments while retaining dense token-wise on-policy distillation supervision; on mathematical reasoning and competitive programming tasks, Seg-OPD-trained students achieve higher revision success rates than baselines and consistently outperform the compared state-of-the-art baselines in reasoning accuracy, with an average relative improvement of 5.22% across diverse models and tasks.