Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

SF-MOPD pairs a fast student with a slow EMA reference to decouple capability interference in multi-teacher distillation, improving multimodal specialization while reducing general-capability degradation across model scales

The work proposes Slow-Fast Multi-Teacher On-Policy Distillation (SF-MOPD), which couples a fast model updated directly by each teacher with a slow model that is an exponential moving average of the student, computes each teacher-induced update in log-probability space, and removes only the component pushing the fast model away from the slow model while retaining aligned and orthogonal components; experiments across multiple model scales show it mitigates capability interference, enhances specialized multimodal capabilities, and reduces average degradation on general-capability benchmarks, consistently outperforming vanilla MOPD.