Skip to main content
Back to timeline
arXivSource publication:

MOPD-Router replaces domain-label hard routing with token-level routing, lifting multi-teacher distillation overall score by 5.88 points on unlabeled mixtures

Synopsis

The work introduces MOPD-Router, a multi-teacher on-policy distillation framework that needs neither domain labels nor a separately trained routing model, selecting and weighting the full teacher pool at every token, and proposes ExpertAlign, a metric that scores each teacher by the positive cosine alignment between the specialization it acquired relative to the shared pre-RL base and the teaching direction it would apply to the current student; across unlabeled and domain-labeled training mixtures under strong-to-weak and same-size distillation, ExpertAlign achieves the strongest overall performance in all four settings, improving the overall score by 5.88 points (+12.3%) over Mean aggregation on unlabeled data and by 3.95 points (+7.

AI-generated editorial illustration: MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation

Interpretation

MOPD-Router changes teacher assignment in multi-teacher on-policy distillation from one domain-matched teacher fixed for an entire rollout per prompt to per-token scoring and weighting over the full teacher pool via plug-in metrics. Existing MOPD practice relies on prompt-level domain labels for hard routing, which restricts the use of unlabeled training mixtures and leaves complementary signals from other teachers unused at individual tokens; this framework uses no domain labels and trains no separate routing model. The paper formalizes sampled-token MOPD, aggregating per-teacher advantages with non-negative weights, and shows standard MOPD and Mean aggregation are special cases of that form; experiments cover unlabeled and labeled data under strong-to-weak and same-size distillation.

ExpertAlign defines each teacher's expertise vector from its policy shift relative to the shared pre-RL base, then aligns it by positive cosine similarity with the teaching vector it would apply to the current student, retaining only positively aligned teachers and weighting them by alignment strength. Entropy looks only at a teacher's own confidence and Novelty only at teacher-student distributional difference, so neither judges whether the correction expresses the specialization the teacher acquired during post-training; ExpertAlign uses the consistency between specialization direction and teaching direction as the routing criterion. The paper reports ExpertAlign as the best overall in all four settings: 38.58 for the 1.7B student and 53.76 for the 4B student on unlabeled data, and 40.19 and 54.58 on labeled data; paired bootstrap tests show it significantly outperforms Mean aggregation and standard MOPD and outperforms or matches Open-MOPD.

Ablations separate ExpertAlign into which teachers contribute and how they are weighted: the positive-alignment gate (Uniform) raises the overall score over Mean aggregation by 1.22-4.75 points across all settings, and cosine weighting adds a further 1.62-3.34 points. This decomposition shows the alignment gate alone selects a more effective subset of teacher signals, while alignment magnitude carries routing information beyond binary selection. The paper tabulates per-domain and overall scores for Mean aggregation, ExpertAlign-Uniform, and ExpertAlign-Cosine in all four settings, noting that Uniform's gains concentrate in mathematics and code while instruction following is slightly lower on unlabeled data.

Cross-domain teachers provide complementary supervision, and that value is best captured through token-level allocation. Restoring the full teacher pool improves the overall score by 2.40 and 2.75 points for the 1.7B and 4B students relative to a domain-restricted variant, and with the pool held constant token-level routing improves over response-level fixed weighting by 3.62 and 1.97 points. The paper reports that all off-domain teachers receive more than 50% of conditional routing mass on math, code, and instruction-following prompts, and gives a case study combining mathematical reasoning, Python verification, and JSON formatting in which the math teacher peaks during numerical derivation, the code teacher during Python computation, and the instruction-following teacher around structural output.

Perspective

The framework targets one specific post-training stage: multi-teacher on-policy distillation where teachers are post-trained from a shared pre-RL base and the student receives dense token-level supervision on its own rollouts. It lets unlabeled SFT and chat mixtures be used directly for multi-teacher distillation without a prior teacher-assignment annotation step; on labeled data it matches or exceeds label-based hard routing without reading the labels. It benefits post-training teams that need to merge mathematics, code, and instruction-following specializations into a single policy model, and researchers who want to compare routing metrics within one MOPD pipeline. The framework code is open-sourced, and the plug-in interface is positioned as a testbed for further metrics incorporating external signals such as verifier feedback or rewards.

The paper itself scopes its conclusions to same-family Qwen3 teachers sharing a pre-RL base, two data regimes, and two distillation scenarios, leaving larger teacher pools and teachers built from different bases as open directions; the authors also describe the routing metrics as guided by empirical intuition rather than a unified design framework. ExpertAlign requires an extra forward pass of the shared base and top-k distribution statistics, measured as a 3.5% increase in GPU-hours in the same-size labeled setting, with that measurement excluding initialization, model downloads, validation, and checkpoint I/O. The top-k support size has a plateau, and at K=32 all three domains degrade because log-ratios amplify differences among near-zero-probability tail tokens; the fraction of fully filtered tokens with no retained teacher is low but shows a mild increasing trend over training, which the paper explains descriptively through a geometric argument. The confidence-interval and three-seed tables in the appendices do not carry concrete numbers in the loaded text, so the magnitude of these uncertainties can only be taken from the ranges stated in the main text.

Sources