Public articles linked to the same research event.
arXiv The work proposes Adaptive Mutual Distillation (AMD), a collaborative post-training framework that jointly trains two models with different task-balancing strategies, evaluates candidate adjustments to distillation weights through short training probes shared across tasks, and uses task-wise validation scores to select an adjustment for each task and transfer direction; across six benchmarks and three LLM backbones both AMD models score higher on average than SFT baselines trained with the same sampling strategies and outperform the task-balancing methods evaluated, and merging the two trained models further improves the average benchmark score while yielding a single model for inference, with merged models exceeding multi-task SFT by an average of 2.91 points across the three backbones.
The work proposes Adaptive Mutual Distillation (AMD), a collaborative post-training framework that jointly trains two models with different task-balancing strategies, evaluates candidate adjustments to distillation weights through short training probes shared across tasks, and uses task-wise validation scores to select an adjustment for each task and transfer direction; across six benchmarks and three LLM backbones both AMD models score higher on average than SFT baselines trained with the same sampling strategies and outperform the task-balancing methods evaluated, and merging the two trained models further improves the average benchmark score while yielding a single model for inference, with merged models exceeding multi-task SFT by an average of 2.91 points across the three backbones.
The work proposes Adaptive Mutual Distillation (AMD), a collaborative post-training framework that jointly trains two models with different task-balancing strategies, evaluates candidate adjustments to distillation weights through short training probes shared across tasks, and uses task-wise validation scores to select an adjustment for each task and transfer direction; across six benchmarks and three LLM backbones both AMD models score higher on average than SFT baselines trained with the same sampling strategies and outperform the task-balancing methods evaluated, and merging the two trained models further improves the average benchmark score while yielding a single model for inference, with merged models exceeding multi-task SFT by an average of 2.91 points across the three backbones.
The work proposes Adaptive Mutual Distillation (AMD), a collaborative post-training framework that jointly trains two models with different task-balancing strategies, evaluates candidate adjustments to distillation weights through short training probes shared across tasks, and uses task-wise validation scores to select an adjustment for each task and transfer direction; across six benchmarks and three LLM backbones both AMD models score higher on average than SFT baselines trained with the same sampling strategies and outperform the task-balancing methods evaluated, and merging the two trained models further improves the average benchmark score while yielding a single model for inference, with merged models exceeding multi-task SFT by an average of 2.91 points across the three backbones.