Skip to main content
Back to timeline
arXivSource publication:

AMD trains two models with different task-balancing strategies and mutual distillation, beating multi-task SFT on six benchmarks and three backbones with a 2.91-point average gain after merging

Related research and updates

Synopsis

The work proposes Adaptive Mutual Distillation (AMD), a collaborative post-training framework that jointly trains two models with different task-balancing strategies, evaluates candidate adjustments to distillation weights through short training probes shared across tasks, and uses task-wise validation scores to select an adjustment for each task and transfer direction; across six benchmarks and three LLM backbones both AMD models score higher on average than SFT baselines trained with the same sampling strategies and outperform the task-balancing methods evaluated, and merging the two trained models further improves the average benchmark score while yielding a single model for inference, with merged models exceeding multi-task SFT by an average of 2.91 points across the three backbones.

Source-provided article image: Adaptive Mutual Distillation for Balanced Multi-Task Post-Training of Large Language Models
Figure 1 ·

Figure 1: Overview of AMD . Two models with different task-balancing strategies undergo independent SFT warmup, followed by mutual distillation with task- and direction-specific weights adjusted using validation feedback. The trained models can optionally be merged into a single model for inference after collaborative training concludes.

arXiv

Interpretation

It introduces Adaptive Mutual Distillation (AMD), a collaborative post-training framework that jointly trains two models using different task-balancing strategies so that each provides cross-model supervision to the other. Existing methods focus primarily on balancing task contributions within single-model training; AMD turns the complementary strengths produced by different balancing strategies into bidirectional distillation signals, moving the balancing problem from inside one model to between models. At the abstract level the framework is specified as joint training of two models, short training probes to evaluate candidate distillation-weight adjustments, and task-wise validation scores to select an adjustment per task and transfer direction; probe length, weight search space, and training steps are not given in the text.

Both AMD models achieve higher average benchmark scores than SFT baselines trained with the same sampling strategies, and they also outperform the task-balancing methods evaluated in the experiments. Relative to single-model task-balancing methods, the gain comes from cross-model, per-task, per-direction weight adaptation rather than from adjusting sampling ratios alone. The evaluation spans six benchmarks and three LLM backbones; the text reports average benchmark-score comparisons without per-benchmark numbers, variance, or significance tests.

Merging the two trained models can further improve their average benchmark score while yielding a single model for inference; the merged models outperform multi-task SFT by an average of 2.91 points across the three backbones. It extends the benefit of mutual distillation from 'each of two models is better' to 'the merged single model is better', reducing the cost of deploying two models. The text explicitly states the 2.91-point average gap and the three-backbone scope; the merging procedure, such as weight interpolation or parameter averaging, is not described in the text.

Perspective

The work targets multi-task post-training where task training-data amounts are unequal and where two models with different task-balancing strategies can be trained jointly; its gains are measured as average benchmark scores over six benchmarks and three LLM backbones, and the merged model targets deployment settings that need a single model for inference. For researchers and engineering teams willing to accept dual-model training and probe-evaluation overhead in exchange for reducing inter-task imbalance, AMD offers a route to convert complementary strengths into distillation signals.

The text is abstract-level and reports no per-benchmark scores, variance, or significance tests, so how the 2.91-point average gain distributes across tasks remains to be seen; the length of the short training probes, how candidate adjustments are searched, how often distillation weights are updated, and how the two models are merged (for example weight interpolation or parameter averaging) are not described in the text; whether task-wise validation-score selection remains stable across different combinations of task-balancing strategies, backbone scales, and numbers of tasks is an open question worth watching.

Sources