Public articles linked to the same research event.
arXiv The work introduces Lexicographic Multi-Objective On-Policy Distillation (LMOPD), which for each student rollout selects the specialist for the first objective whose gate detects a deficiency and locally projects its centered log-policy correction to remove components opposing higher-priority specialists; on three math benchmarks with 30B-A3B mixture-of-experts models, the two-expert setting fully retains accuracy and reasoning-quality gains in point estimates while acquiring 46.9% of the conciseness gain, and the four-expert setting retains about 90% of both the accuracy gain and reasoning-correctness gain versus only about 57% for the next best evaluated baseline.
The work introduces Lexicographic Multi-Objective On-Policy Distillation (LMOPD), which for each student rollout selects the specialist for the first objective whose gate detects a deficiency and locally projects its centered log-policy correction to remove components opposing higher-priority specialists; on three math benchmarks with 30B-A3B mixture-of-experts models, the two-expert setting fully retains accuracy and reasoning-quality gains in point estimates while acquiring 46.9% of the conciseness gain, and the four-expert setting retains about 90% of both the accuracy gain and reasoning-correctness gain versus only about 57% for the next best evaluated baseline.
The work introduces Lexicographic Multi-Objective On-Policy Distillation (LMOPD), which for each student rollout selects the specialist for the first objective whose gate detects a deficiency and locally projects its centered log-policy correction to remove components opposing higher-priority specialists; on three math benchmarks with 30B-A3B mixture-of-experts models, the two-expert setting fully retains accuracy and reasoning-quality gains in point estimates while acquiring 46.9% of the conciseness gain, and the four-expert setting retains about 90% of both the accuracy gain and reasoning-correctness gain versus only about 57% for the next best evaluated baseline.
The work introduces Lexicographic Multi-Objective On-Policy Distillation (LMOPD), which for each student rollout selects the specialist for the first objective whose gate detects a deficiency and locally projects its centered log-policy correction to remove components opposing higher-priority specialists; on three math benchmarks with 30B-A3B mixture-of-experts models, the two-expert setting fully retains accuracy and reasoning-quality gains in point estimates while acquiring 46.9% of the conciseness gain, and the four-expert setting retains about 90% of both the accuracy gain and reasoning-correctness gain versus only about 57% for the next best evaluated baseline.