Skip to main content
Back to timeline
arXivSource publication:

OG-OPSD selects KL direction by outcome correctness and truncates prefixes by entropy, lifting Qwen3 1.7B/4B/8B math averages to 44.4/64.9/66.4

Related research and updates

Synopsis

Analyzing on-policy self-distillation (OPSD) through the advantage formulation of RLVR, this work finds that a fixed divergence objective under-penalizes and over-rewards incorrect trajectories and that teacher supervision reliability varies with trajectory outcome and cumulative average teacher entropy; it proposes OG-OPSD, which applies forward KL across correct trajectories and reverse KL only within a prefix of incorrect trajectories whose cutoff is set by the last crossing of the cumulative average teacher entropy curves, consistently outperforming vanilla OPSD and multiple strong baselines on mathematical reasoning, multimodal reasoning, and out-of-distribution tasks across Qwen3-1.7B/4B/8B and Qwen3-VL-2B.

Source-provided article image: Outcome-Guided On-Policy Self-Distillation

(b) Qwen3-4B

arXiv

Interpretation

It links the implicit token-level advantages of forward KL and reverse KL to trajectory correctness: the two share the same sign but differ in coefficient magnitude, with FKL giving larger positive advantages to teacher-preferred tokens and a bounded negative penalty, while RKL gives unbounded negative penalties to student-overweighted tokens. Vanilla OPSD applies a fixed divergence across all trajectories and positions, and recent work relies on token-level statistics or extra hyperparameters; this work instead uses the binary verifier outcome as a trajectory-level selection signal. Gradient derivations in the REINFORCE form (Appendix A) yield the implicit advantages in Eqs. 6, 7, and 8, and ablations comparing fixed FKL, fixed RKL, and reversed pairing all show performance drops.

It identifies an outcome-dependent prefix-failure pattern: on incorrect trajectories, teacher completion accuracy degrades markedly as the truncated prefix grows longer, whereas on correct trajectories it remains stable or even improves. It moves the account of declining teacher reliability from the token-position level to the trajectory-outcome level, showing that truncation should not be applied uniformly to all trajectories. Responses sampled from Qwen3 models on OpenThoughts-Math30K, DeepScaleR, MATH-500, and DeepMath are split into correct and incorrect pools by verifier outcome, and teacher completion accuracy (TCA) is measured at varying truncation lengths, with the same trend observed on Qwen3-1.7B.

It proposes OG-OPSD: full FKL on correct trajectories, and RKL on incorrect trajectories only within a prefix set by the last crossing of the cumulative average teacher entropy curves for the two outcome groups, with that cutoff shared across incorrect trajectories. Compared with uniform truncation in ESR, position weighting in PW-OPSD, and token-level entropy-based divergence selection in EOPD, this method uses only the binary outcome and cumulative average entropy, adding no extra hyperparameters or trade-offs. Ablations show consistent drops when truncation is removed, applied to all trajectories, or when the divergence pairing is reversed; performance is relatively stable across neighboring values of the truncation parameter.

On Qwen3-1.7B/4B/8B, OG-OPSD reaches three-benchmark averages of 44.4, 64.9, and 66.4 on AIME24, AIME25, and HMMT25, above vanilla OPSD and the listed baselines; on Qwen3-VL-2B it also attains the highest scores among evaluated methods on FinMME and on MathVerse, MathVision, and MMMUPro. The gain over vanilla OPSD widens with model scale, and the drop on LiveCodeBench v5 after mathematical post-training is smaller than for vanilla OPSD and PW-OPSD. Mathematical evaluation uses Avg@12 over 30 problems each on AIME24/AIME25/HMMT25; multimodal evaluation uses the FinMME test set plus three OOD benchmarks; code retention is measured on LCB v5.

Perspective

The results target task settings with automatically verifiable binary outcomes, such as mathematical reasoning and multimodal mathematical reasoning; the method is validated on Qwen3-1.7B/4B/8B Instruct and Qwen3-VL-2B-Instruct, trained on OpenThoughts-Math30K (29,434 examples) and a FinMME subset, with the cutoff determined on training data from the last crossing of the cumulative average teacher entropy curves for correct and incorrect groups and shared across incorrect trajectories. For teams seeking to reduce self-distillation noise without adding hyperparameters, this rule of choosing divergence by outcome and position by entropy can be embedded directly into existing OPSD pipelines; the authors suggest future extension to graded rewards and per-trajectory adaptive truncation.

Teacher completion accuracy and cumulative average entropy curves are presented in Figures 1 and 4 without specific numerical values in the main text, so the cutoff value under different data distributions needs independent reproduction; the sensitivity analysis shows relative stability near the peak, but the applicable range at larger scale or in non-mathematical domains remains to be tested; the authors note that compute limits prevented comparing all baselines at 1.7B and 8B, so the cross-scale picture is not fully controlled; multimodal and code-retention results come from a single VL model and LCB v5, and their transferability deserves further observation.

Sources