Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Latent-MOPD lets one student absorb both the predictions and hidden states of multiple specialists, beating single-channel baselines on all nine math, code, and logic benchmarks

The work introduces Latent-MOPD, a multi-teacher on-policy distillation method for LLMs that, to the authors' knowledge, is the first to integrate specialists at the representation level: it selects late-layer hidden-state targets according to the teacher-student relationship, bridges unequal hidden widths with a shared projection, groups updates by domain, and gradually shifts each teacher's supervision from hidden states to token predictions; in the same-family setting it outperforms token-only, representation-only, and uniform-averaging baselines on all nine benchmarks across math, code, and logic, surpasses the per-benchmark best teacher on a majority of them at equal parameter count, and with larger cross-family teachers it outperforms both single-channel baselines on all benchmarks.