Skip to main content
Back to timeline
arXivSource publication:

DN-MOPD normalizes per-domain teacher feedback and lifts the six-task average over MOPD from 58.4/50.3/26.6 to 59.6/52.5/29.0 across three Qwen3.5 sizes

Synopsis

The work finds that multi-teacher on-policy distillation (MOPD) with domain-label routing alone fails to beat the strongest single-teacher student and transfers little of the mathematics expert's gain, because domain feedback is unbalanced (in the first training batch instruction-following log-ratios are 2.3-4.4 times as dispersed as the pooled signal while mathematics is about half as dispersed, and for the initial 4B student the instruction-following loss supplies 94% of the combined gradient), and proposes DN-MOPD, which keeps label routing and rescales each domain's distillation advantages by the clipped ratio of pooled to domain log-ratio spread estimated every batch, improving the six-task average over MOPD at every size and both evaluation budgets (1.17-2.36 points at 16K, 2.47-3.

AI-generated editorial illustration: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation

Interpretation

Teacher assignment alone does not transfer specialist skills: on Qwen3.5-9B/4B/2B, label-routed MOPD never outperforms the strongest single-teacher student at any size and transfers little of the mathematics expert's gain (retaining only 14-32% at the 8K budget and showing no gain over initialization at 16K). Prior MOPD work and its follow-ups focused on which expert teaches (routing) or set weights by hand from token share and the remaining teacher-student gap; this work isolates feedback strength as a separate design dimension and quantifies its imbalance. Three independently trained Qwen3.5 expert pools, shared experts, student initialization and prompts within each pool, 80 updates, student seed 42 (plus two more seeds), with paired 95% intervals resampling questions; the imbalance is supported by first-batch log-ratio dispersion and per-domain gradients of the initial 4B student.

DN-MOPD puts every teacher's feedback on a common scale: it rescales each domain's distillation advantages by the clipped ratio of pooled to domain log-ratio standard deviation estimated on every batch, keeping label routing and the sign of every advantage, with MOPD as the special case where every multiplier equals one. Unlike weighting by token share or mean reward magnitude (as in Open-MOPD's domain budgets and reward refresh), the rule targets dispersion rather than teacher quality or remaining capability gap, and needs no extra teacher call, teacher model or learned router. The rule is fully specified in Section 2 and App. D.1 (population standard deviation convention, clipping bounds, the degenerate multiplier-one case) and released with a reference implementation and unit tests; measured multipliers are about 1.5-2.1 for mathematics, near 1 for code at 4B/2B, and at the 0.25 floor for instruction following in most batches.

Calibrating feedback scale consistently improves on MOPD: DN-MOPD improves the six-task average at every size and both evaluation budgets with all paired intervals above zero, holds across three student seeds, exceeds the strongest single-teacher student, and recovers most of the mathematics gain MOPD loses. Relative to label-routed MOPD it gains 1.17-2.36 points at 16K and 2.47-3.08 at 8K; mathematics improves at every size, budget and training endpoint with intervals above zero, while code and instruction-following differences are smaller and mixed. The main comparison holds experts, student initialization, prompts and update budget fixed; seed-matched gains average +1.12, +1.97 and +2.34 at 16K with two-level intervals excluding zero; on MATH-500 DN-MOPD leads Label at every size and cap with intervals above zero.

The gain comes mainly from turning down instruction-following feedback rather than amplifying mathematics alone: fixed-weight controls show that reducing IF to 0.25 alone recovers most of the improvement at 4B and 2B and raises mathematics by about three points, while amplifying mathematics to 2 helps less, and fixed weights near DN-MOPD's measured multipliers perform comparably at 9B and 4B. This explains why the intuitive remedy of amplifying the under-transferred mathematics domain recovers only about half of the improvement at 4B and little at 2B, and separates feedback-scale calibration from teacher selection. Fixed-weight controls cover 9B/4B/2B and both evaluation caps, including update-0 weights, a global (2,1,0.25) setting and single-domain controls; per-domain gradient measurements for the initial 4B student show IF supplying 94% of the combined gradient under equal weights and 64% under DN-MOPD's first-batch weights.

Perspective

The result targets expert pools for mathematics, code and instruction following trained by separate RL pipelines within one model family, integrated into a shared student; the method reuses log-ratios already available from rollout scoring and adds no teacher call, teacher model or learned router, so it suits teams with an existing MOPD pipeline who want to adjust feedback strength without changing routing or data mix. It applies to settings with domain-label routing and per-batch within-domain log-ratio statistics, and was validated on Qwen3.5-9B/4B/2B at both 8K and 16K evaluation budgets; at 2B, per-batch re-estimation adds about 1.10 points over fixed first-batch weights.

Comparisons use one expert pool per size within one model family, and the three student seeds do not cover retraining the experts; the method was selected on an earlier Qwen3 development instrument and an earlier Qwen3-4B construction found no clear gain, so benefits depend on the teacher-student configuration. Scale estimates depend on domain composition and response length, the clipping bounds were not tuned per size, and the instruction-following multiplier usually sits at the lower bound, so the rule bounds rather than equalizes that domain's scale. Fixed weights perform comparably to DN-MOPD at 9B and 4B, with per-batch re-estimation adding only at 2B. Longer evaluation budgets narrow the gap to label routing, and the eventual ordering at convergence after 160 updates remains unresolved; SeqKD-SFT and task arithmetic retain higher Totals under their own recipes. The per-domain gradient measurements use one probe set, policy ratio one without clipping or optimizer state, and are descriptive. In addition, several appendix tables in the loaded text have empty numeric cells, so paired intervals and some per-domain contrasts can only be read from the prose rather than checked number by number.

Sources