Latent-MOPD lets one student absorb both the predictions and hidden states of multiple specialists, beating single-channel baselines on all nine math, code, and logic benchmarks
Related research and updatesSynopsis
The work introduces Latent-MOPD, a multi-teacher on-policy distillation method for LLMs that, to the authors' knowledge, is the first to integrate specialists at the representation level: it selects late-layer hidden-state targets according to the teacher-student relationship, bridges unequal hidden widths with a shared projection, groups updates by domain, and gradually shifts each teacher's supervision from hidden states to token predictions; in the same-family setting it outperforms token-only, representation-only, and uniform-averaging baselines on all nine benchmarks across math, code, and logic, surpasses the per-benchmark best teacher on a majority of them at equal parameter count, and with larger cross-family teachers it outperforms both single-channel baselines on all benchmarks.
Figure 1: Same-family gains beyond individual teachers, with a cross-family extension. (a) Same-family; (b) cross-family (last 1, Linear). Both panels use shared per-axis scales, with each outer vertex fixed at the same-family Latent-MOPD score. Solid polygons show Latent-MOPD; the dotted ring in (b) repeats the same-family reference. Representation-only denotes OPRD-style (all layers) in (a) and Rep-only (last 1) in (b) . Labels give Latent-MOPD scores; bold with ( > > Teacher) marks those above the strongest teacher. Table 1 reports all nine same-family benchmarks.
arXivInterpretation
It proposes Latent-MOPD, described as the first representation-level multi-teacher on-policy distillation method for LLMs, integrating specialists through both their predictions and the hidden states used to compute them, without additional teacher training. Prior multi-teacher on-policy distillation transferred only what specialists predict through their output distributions; this work extends supervision to internal hidden states so a single student uses both channels. The abstract states the method-level positioning and design, supported by same-family and cross-family experiments.
To coordinate representation supervision from multiple specialists, the method selects late-layer targets according to the teacher-student relationship, bridges unequal hidden widths with a shared projection, and groups updates by domain; each teacher's supervision gradually shifts from hidden states to token predictions, with both channels using the same routed specialist. These mechanisms address layer selection, width mismatch, and domain mixing in multi-teacher representation supervision rather than simply averaging teacher signals. The abstract lists these mechanisms as design elements and notes that both channels share the same routed specialist.
In the main same-family setting, Latent-MOPD outperforms the token-only, representation-only, and uniform-averaging baselines on all nine benchmarks across math, code, and logic; with the same parameter count as each teacher, the student also surpasses the per-benchmark best teacher on a majority of these benchmarks. Relative to existing single-channel and uniform-averaging approaches, the method shows an advantage across all nine benchmarks and indicates a student can exceed individual specialist teachers. The abstract reports comparisons on nine benchmarks and states the student matches each teacher's parameter count.
With larger, separately developed cross-family teachers, Latent-MOPD outperforms both single-channel baselines on all benchmarks; a same-family all-layer representation-only control remains stable with domain-pure updates but collapses when teacher domains are interleaved within an update. The cross-family results indicate the method is not limited to same-family teachers, while the control reveals that domain interleaving strongly affects representation-only supervision stability. The abstract gives the cross-family all-benchmark comparison and the stability difference of the same-family all-layer representation-only control under two update organizations.
Perspective
The work targets multi-teacher on-policy distillation for LLMs in settings where expert hidden states are accessible and supervision layers can be chosen from the teacher-student relationship. The abstract indicates its value in the same-family setting, where it outperforms token-only, representation-only, and uniform-averaging baselines on all nine benchmarks across math, code, and logic and surpasses the per-benchmark best teacher on a majority at equal parameter count, and in the cross-family setting, where it outperforms both single-channel baselines on all benchmarks. It therefore mainly serves research and engineering practice that seeks to integrate several domain specialists into one student, especially when teachers are same-family or are larger and separately developed. The abstract does not give benchmark names, model scales, training data, or hyperparameters, so reproduction should follow the original paper and accompanying code.
The abstract does not list the nine benchmark names, teacher and student model scales, training data composition, or hyperparameters, nor does it give per-benchmark numbers or effect sizes, so the magnitude of the advantages is hard to judge. The same-family all-layer representation-only control collapsing under interleaved-domain updates suggests update organization affects representation-supervision stability, but the mechanism and scope remain to be confirmed in the paper. The cross-family results state an advantage over both single-channel baselines but do not compare against each cross-family teacher individually. These are information boundaries at the abstract level; specific conclusions should be checked against the paper's main text, figures, and supplementary materials.
