Public articles linked to the same research event.
arXiv Analyzing on-policy self-distillation (OPSD) through the advantage formulation of RLVR, this work finds that a fixed divergence objective under-penalizes and over-rewards incorrect trajectories and that teacher supervision reliability varies with trajectory outcome and cumulative average teacher entropy; it proposes OG-OPSD, which applies forward KL across correct trajectories and reverse KL only within a prefix of incorrect trajectories whose cutoff is set by the last crossing of the cumulative average teacher entropy curves, consistently outperforming vanilla OPSD and multiple strong baselines on mathematical reasoning, multimodal reasoning, and out-of-distribution tasks across Qwen3-1.7B/4B/8B and Qwen3-VL-2B.
Analyzing on-policy self-distillation (OPSD) through the advantage formulation of RLVR, this work finds that a fixed divergence objective under-penalizes and over-rewards incorrect trajectories and that teacher supervision reliability varies with trajectory outcome and cumulative average teacher entropy; it proposes OG-OPSD, which applies forward KL across correct trajectories and reverse KL only within a prefix of incorrect trajectories whose cutoff is set by the last crossing of the cumulative average teacher entropy curves, consistently outperforming vanilla OPSD and multiple strong baselines on mathematical reasoning, multimodal reasoning, and out-of-distribution tasks across Qwen3-1.7B/4B/8B and Qwen3-VL-2B.
Analyzing on-policy self-distillation (OPSD) through the advantage formulation of RLVR, this work finds that a fixed divergence objective under-penalizes and over-rewards incorrect trajectories and that teacher supervision reliability varies with trajectory outcome and cumulative average teacher entropy; it proposes OG-OPSD, which applies forward KL across correct trajectories and reverse KL only within a prefix of incorrect trajectories whose cutoff is set by the last crossing of the cumulative average teacher entropy curves, consistently outperforming vanilla OPSD and multiple strong baselines on mathematical reasoning, multimodal reasoning, and out-of-distribution tasks across Qwen3-1.7B/4B/8B and Qwen3-VL-2B.
Analyzing on-policy self-distillation (OPSD) through the advantage formulation of RLVR, this work finds that a fixed divergence objective under-penalizes and over-rewards incorrect trajectories and that teacher supervision reliability varies with trajectory outcome and cumulative average teacher entropy; it proposes OG-OPSD, which applies forward KL across correct trajectories and reverse KL only within a prefix of incorrect trajectories whose cutoff is set by the last crossing of the cumulative average teacher entropy curves, consistently outperforming vanilla OPSD and multiple strong baselines on mathematical reasoning, multimodal reasoning, and out-of-distribution tasks across Qwen3-1.7B/4B/8B and Qwen3-VL-2B.