Skip to main content
Back to timeline
arXivSource publication:

PMOPD projects parameter updates away from protected task subspaces to ease the multi-teacher distillation capability seesaw, lifting the three-task average by 2.54 points on Qwen2.5-7B and 2.09 on Llama-3.1-8B

Synopsis

The work proposes PMOPD, which builds per-matrix low-dimensional subspace memories from the cumulative parameter displacements of completed task blocks and projects both gradients and Adafactor optimizer updates of later tasks to remove components conflicting with protected task directions, while a lightweight conflict probe sets task order and a subspace-consistency criterion sets the cycle count; across Code, Reason, and Math, PMOPD improves every evaluated capability over MOPD, raising the three-task average by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B.

AI-generated editorial illustration: PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation

Interpretation

The paper characterizes cross-task interference in multi-teacher on-policy distillation (MOPD) geometrically: cumulative parameter updates from different tasks rapidly concentrate in their own low-dimensional subspaces, consistent within a task and with low cross-task overlap. Prior OPD work largely optimized single-task distillation through objective design, distillation scope, and teacher signal construction, whereas this work turns the MOPD capability seesaw into a measurable subspace-overlap problem. The text reports that subspaces extracted after 20% of training already reach an average similarity of 0.62 to their final subspaces, and that similarities among the top-16 update subspaces for Math, Code, and Reason range from 0.133 to 0.151.

PMOPD builds per-matrix subspace memories from the realized displacements of completed task blocks and projects both the gradient and the preconditioned optimizer update of later tasks to remove components aligned with protected directions. Unlike multi-task methods such as PCGrad that act on instantaneous gradients, PMOPD protects realized OPD block displacements that include clipping and adaptive-optimization effects; the text notes that projecting the gradient alone is insufficient because Adafactor's element-wise adaptive preconditioning can rotate an already projected gradient back toward protected subspaces. Under the same task order, four cycles, and identical data budget, the ablation moves the average from 63.85 with no projection to 66.22 with gradient projection and to 66.97 with optimizer-update projection added, a 3.12-point gain over the matched no-projection variant.

The paper introduces a lightweight conflict probe and a cycling strategy to organize task updates, ranking tasks by bidirectional conflict measurements and choosing the cycle count by same-task cross-cycle subspace similarity. This converts two scheduling choices, task order and cycle count, from exhaustive search into low-cost geometric diagnostics that can be run before full training. The six task orders span 2.97 points in average score, with CodeReasonMath highest at 65.81; all ten folds of 60 examples each recover the Code, Reason, Math ordering, with means of 28.04%, 35.75%, and 36.97%; the cycle sweep peaks at four cycles with the highest average of 66.97 and the highest cross-cycle similarity of 0.525.

Across two model families and three capability domains, PMOPD improves every evaluated capability over MOPD rather than redistributing performance among domains. The gains appear on both Qwen2.5-7B and Llama-3.1-8B, indicating the geometry-aware optimization rule is not tied to a single model configuration. On Qwen2.5-7B the average is 66.97, 2.54 points above MOPD and 3.08 above BB-MOPD; on Llama-3.1-8B the average is 41.04, 2.09 points above MOPD and 4.75 above BB-MOPD; PMOPD results are averaged over three independent runs with different random seeds.

Perspective

The results target post-training settings that integrate multiple domain teachers into shared parameters through on-policy distillation, with teachers derived by domain-specific reinforcement learning from the same base checkpoint, and are validated on Math, Reason, and Code across Qwen2.5-7B and Llama-3.1-8B. The method builds subspace memories independently for every two-dimensional trainable parameter matrix and rebuilds memory from an empty state each cycle, so the protected geometry tracks the current trajectory without an unbounded archive. For practitioners who want to set task order and cycle count without exhaustive permutation search, the conflict probe and cross-cycle subspace similarity provide pre-training selection signals.

Conclusions about task order and cycle count rest on three tasks, two model families, and a specific data scale, so whether the conflict ranking and best cycle count persist as the number of tasks or the training scale grows remains open. The conflict probe recovers the ordering stably with 60 examples per fold, but how probe sample size affects ranking stability is not explored in the text. Four cycles peak simultaneously in average score and cross-cycle similarity, which supports the mechanistic account, yet the causal chain between the two would need more settings to test. This reading covered the full text, including the appendix per-task similarities and directional conflict measurements; unseen supplementary material may hold further detail.

Sources