Skip to main content
Back to timeline
arXivSource publication:

SF-MOPD pairs a fast student with a slow EMA reference to decouple capability interference in multi-teacher distillation, improving multimodal specialization while reducing general-capability degradation across model scales

Related research and updates

Synopsis

The work proposes Slow-Fast Multi-Teacher On-Policy Distillation (SF-MOPD), which couples a fast model updated directly by each teacher with a slow model that is an exponential moving average of the student, computes each teacher-induced update in log-probability space, and removes only the component pushing the fast model away from the slow model while retaining aligned and orthogonal components; experiments across multiple model scales show it mitigates capability interference, enhances specialized multimodal capabilities, and reduces average degradation on general-capability benchmarks, consistently outperforming vanilla MOPD.

Source-provided article image: Slow-Fast Multi-Teacher On-Policy Distillation for Capability Preservation
Figure 1 ·

Figure 1 : Comparison of capability absorption by MOPD and SF-MOPD across 11 benchmarks, with scores normalized separately for each dataset. The green line denotes the per-benchmark oracle reference obtained from the corresponding expert. SF-MOPD approaches the oracle reference on most benchmarks while better preserving general capabilities.

arXiv

Interpretation

The paper states that multi-teacher on-policy distillation (MOPD) gradually drives the student away from its initialization model when consolidating domain expertise into a single student, and general capabilities decline as this displacement grows, producing capability interference. Where MOPD was described as an effective framework for consolidating domain expertise, this side effect is named as a problem to be addressed rather than a training detail. This claim comes from the abstract's description of MOPD training dynamics; it is a problem-formulation statement without quantitative curves in the abstract.

The paper proposes SF-MOPD: a fast model that is the current student updated directly by each teacher, paired with a slow model that is an exponential moving average of the student, absorbing the learning signal gradually and serving as a moving capability reference that fuses the general foundation with confirmed domain expertise. Unlike directly constraining the student toward its initialization, the slow model is not a fixed anchor but an evolving reference, so domain expertise acquisition is not suppressed. The method description comes from the abstract, including the fast/slow definitions and the slow model as an exponential moving average; no hyperparameters or update equations are given in the abstract.

For each teacher, SF-MOPD computes the teacher-induced update in log-probability space and removes only the component that pushes the fast model further away from the slow model, retaining aligned and orthogonal components. Capability preservation is turned into a decomposition of the update direction with selective removal, rather than an overall suppression of student drift. The abstract states this decomposition and retention rule explicitly, but provides no ablation or component-contribution numbers.

Experiments across multiple model scales show SF-MOPD mitigates capability interference, enhances specialized multimodal capabilities, and reduces average degradation on general-capability benchmarks, consistently outperforming vanilla MOPD. Relative to vanilla MOPD, the reported direction is simultaneous improvement in specialization and general capability rather than trading one for the other. Evidence comes from the multi-scale experiments and benchmark comparison described in the abstract; specific benchmark names, model scales, and values are not listed.

Perspective

The work targets multi-teacher on-policy distillation settings that consolidate several domain teachers into a single student, and it applies to multimodal foundation model training pipelines that aim to retain general foundation capabilities alongside domain expertise. Its mechanism relies on pairing a fast and a slow model and on directional decomposition in log-probability space, so it is most directly usable by teams already using a similar on-policy distillation paradigm; the multi-scale experiments described in the abstract suggest the idea can be tried at different scales.

The abstract does not state the specific benchmarks, model scales, number of teachers, or training cost, nor the effect of the slow model's update rate; the trade-off between the magnitude of interference mitigation and the magnitude of specialization gain still needs confirmation in the main experiments. In addition, this summary is based only on the abstract, without the main text, figures, or appendix, so the full evidence chain for the mechanism details and experimental conclusions remains an open question to verify.

Sources