Skip to main content
Back to timeline
arXivSource publication:

LMOPD uses lexicographic multi-teacher distillation to preserve top-priority capabilities at about 90% retention on a 30B-A3B model

Related research and updates

Synopsis

The work introduces Lexicographic Multi-Objective On-Policy Distillation (LMOPD), which for each student rollout selects the specialist for the first objective whose gate detects a deficiency and locally projects its centered log-policy correction to remove components opposing higher-priority specialists; on three math benchmarks with 30B-A3B mixture-of-experts models, the two-expert setting fully retains accuracy and reasoning-quality gains in point estimates while acquiring 46.9% of the conciseness gain, and the four-expert setting retains about 90% of both the accuracy gain and reasoning-correctness gain versus only about 57% for the next best evaluated baseline.

Source-provided article image: Lexicographic Multi-Objective On-Policy Distillation
Figure 1 ·

Figure 1: LMOPD overview. (a) We train reward-specialized experts and assign each reward to an expert; one expert may serve multiple rewards. (b) Each student rollout is checked in that order and routed to the first objective whose gate detects a deficiency; a rollout that passes every gate is self-distilled. (c) At each token, LMOPD removes only components of the routed specialist’s correction that oppose earlier specialists, producing a projected virtual-teacher target.

arXiv

Interpretation

Introduces LMOPD, a multi-teacher method that integrates reward-specialized policies under explicit priorities by selecting, for each student rollout, the specialist for the first objective whose gate detects a deficiency, then locally projecting its centered log-policy correction to remove components that oppose higher-priority specialists. Existing multi-reward post-training methods typically scalarize rewards or combine specialists without explicitly protecting a reward priority order; LMOPD makes the priority order central to the integration mechanism rather than an after-the-fact weighting. The method description comes from the abstract, and matched four-expert ablations show lexicographic routing outperforms random routing and that projection further strengthens both top-priority capabilities.

On 30B-A3B mixture-of-experts transformers, the two-expert setting fully retains the accuracy and reasoning-quality gains in point estimates while acquiring 46.9% of the conciseness gain. Shows that under asymmetric trade-offs, conciseness can be partially acquired without sacrificing higher-priority correctness, rather than trading correctness for conciseness. Based on retained-gain measurements on three math benchmarks, reported as point estimates.

In the four-expert setting, LMOPD retains about 90% of both the accuracy gain and the reasoning-correctness gain, compared with only about 57% for the next best evaluated baseline. As the number of experts grows and objective conflicts become more complex, the retention advantage from explicit priorities becomes more pronounced relative to baselines. Retained-gain comparison on three math benchmarks, with baselines as evaluated by the authors.

Matched four-expert ablations indicate that lexicographic routing outperforms random routing and that the projection step further strengthens both top-priority capabilities. Decomposes the overall benefit into routing and projection components, giving the direction of each component's independent contribution. Matched four-expert ablation experiments.

Perspective

The result targets post-training settings that need explicit priority ordering among multiple reward objectives, especially when correctness, reasoning quality, and conciseness involve asymmetric trade-offs. The method applies to integrating multi-teacher, reward-specialized policies and is validated on 30B-A3B mixture-of-experts transformers across three math benchmarks, with quantified retained gains in two- and four-expert settings. For engineering practice seeking secondary-objective gains without sacrificing top-priority capabilities, LMOPD offers a reusable routing-plus-projection approach.

The current evidence is at the abstract level and does not include benchmark names, evaluation details, training costs, or statistical uncertainty, so the robustness of retained gains beyond point estimates and behavior on other tasks and model scales remain open questions. The separate contribution magnitudes of routing and projection in the four-expert ablations, and how results change when the priority order changes, also need confirmation in the full text.

Sources