Public articles linked to the same research event.
arXiv The work proposes Slow-Fast Multi-Teacher On-Policy Distillation (SF-MOPD), which couples a fast model updated directly by each teacher with a slow model that is an exponential moving average of the student, computes each teacher-induced update in log-probability space, and removes only the component pushing the fast model away from the slow model while retaining aligned and orthogonal components; experiments across multiple model scales show it mitigates capability interference, enhances specialized multimodal capabilities, and reduces average degradation on general-capability benchmarks, consistently outperforming vanilla MOPD.
The work proposes Slow-Fast Multi-Teacher On-Policy Distillation (SF-MOPD), which couples a fast model updated directly by each teacher with a slow model that is an exponential moving average of the student, computes each teacher-induced update in log-probability space, and removes only the component pushing the fast model away from the slow model while retaining aligned and orthogonal components; experiments across multiple model scales show it mitigates capability interference, enhances specialized multimodal capabilities, and reduces average degradation on general-capability benchmarks, consistently outperforming vanilla MOPD.
The work proposes Slow-Fast Multi-Teacher On-Policy Distillation (SF-MOPD), which couples a fast model updated directly by each teacher with a slow model that is an exponential moving average of the student, computes each teacher-induced update in log-probability space, and removes only the component pushing the fast model away from the slow model while retaining aligned and orthogonal components; experiments across multiple model scales show it mitigates capability interference, enhances specialized multimodal capabilities, and reduces average degradation on general-capability benchmarks, consistently outperforming vanilla MOPD.
The work proposes Slow-Fast Multi-Teacher On-Policy Distillation (SF-MOPD), which couples a fast model updated directly by each teacher with a slow model that is an exponential moving average of the student, computes each teacher-induced update in log-probability space, and removes only the component pushing the fast model away from the slow model while retaining aligned and orthogonal components; experiments across multiple model scales show it mitigates capability interference, enhances specialized multimodal capabilities, and reduces average degradation on general-capability benchmarks, consistently outperforming vanilla MOPD.