Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

OG-OPSD selects KL direction by outcome correctness and truncates prefixes by entropy, lifting Qwen3 1.7B/4B/8B math averages to 44.4/64.9/66.4

Analyzing on-policy self-distillation (OPSD) through the advantage formulation of RLVR, this work finds that a fixed divergence objective under-penalizes and over-rewards incorrect trajectories and that teacher supervision reliability varies with trajectory outcome and cumulative average teacher entropy; it proposes OG-OPSD, which applies forward KL across correct trajectories and reverse KL only within a prefix of incorrect trajectories whose cutoff is set by the last crossing of the cumulative average teacher entropy curves, consistently outperforming vanilla OPSD and multiple strong baselines on mathematical reasoning, multimodal reasoning, and out-of-distribution tasks across Qwen3-1.7B/4B/8B and Qwen3-VL-2B.