Skip to main content
Back to timeline
arXivSource publication:

DCSD decouples credit direction from magnitude in self-distillation, topping 11 benchmarks and lifting math reasoning by 8.45 points over base models

Synopsis

The work introduces Decoupled Credit Self-Distillation (DCSD), which uses belief-margin probing to set each reasoning step's credit direction and marginal information gain to quantify its contribution magnitude, leaving the privileged teacher to allocate credit only within steps; across 11 mathematical and multimodal reasoning benchmarks DCSD achieves the best overall scores against GRPO, OPSD, RLSD and RLCSD, improving overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning over base models, while correcting credit direction for about 6% of tokens and reducing token credit magnitude by roughly 1.5 times.

AI-generated editorial illustration: Can We Trust the Teacher? Decoupled Credit Direction-Magnitude for Self-Distillation

Interpretation

The paper formalizes "oracle step credit" in self-distillation and shows that under a policy conditioned on eventual success, the self-distillation log-ratio is directionally consistent with the oracle local RL advantage, while its magnitude does not directly represent the step's contribution. Prior OPSD coupled credit direction and magnitude in a single teacher signal, and RLSD/RLCSD anchored direction to the trajectory-level advantage; this work separates "directional consistency" from "magnitude is not directly recoverable" as two properties that can be handled independently. The conclusion comes from the theoretical derivation in Section 2 (Eqs. 1-3 and Appendix B); it is an analytical argument rather than an experimental measurement.

DCSD uses belief-margin probing to decide whether a step shifts the student's belief toward or away from the correct answer, and marginal information gain to quantify how much non-redundant information the step adds beyond its preceding context, thereby fixing credit direction and magnitude separately. Unlike directly using the teacher signal or anchoring direction to the trajectory-level advantage, DCSD restricts the teacher to within-step token allocation, so teacher variation can redistribute credit inside a step but cannot change the established step direction or total magnitude. The method is fully specified in Section 3 and Appendix C, supported by Theorems 1-4; local direction is used only when it crosses a confidence threshold, otherwise the trajectory-level direction is used as fallback.

Across 11 benchmarks DCSD achieves the best overall scores: mathematical reasoning (Qwen3-4B) reaches 88.56 overall versus GRPO 84.70, RLSD 83.86, OPSD 80.11 and vanilla 81.40; multimodal reasoning (Qwen3-VL-8B-Instruct) reaches 61.14 overall versus RLSD 58.19, GRPO 55.49, OPSD 54.13 and vanilla 53.43. Relative to base models, DCSD improves overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning; the largest multimodal gain is on WeMath (54.10 to 67.05), while ZeroBench is the exception where DCSD underperforms RLSD. Results are mean@4 accuracy; mathematical and multimodal baselines follow shared settings and the published RLSD settings respectively, and for multimodal tasks only DCSD is trained while baseline results are taken from Table 2 of RLSD.

Training dynamics and teacher-signal diagnostics show that DCSD corrects credit direction for about 6% of tokens (including negative corrections in successful trajectories and positive corrections in unsuccessful ones) with stable credit-magnitude attenuation during training; on 90 fixed AIME24-26 problems under C0-C4 conditions that vary only the privileged information, DCSD attains lower Top-1 disagreement rate and mean absolute error than OPSD and RLSD. The paper uses a condition-matched Qwen3-32B as an operational oracle reference and notes that DCSD's advantage already appears under C0, where no privileged answer information is provided, so the improvement cannot be attributed solely to additional answer information. Diagnostics use frozen Qwen3-4B responses (1,098,216 tokens per scoring condition) with paired bootstrap intervals; the authors explicitly treat Qwen3-32B as a proxy for the oracle rather than the true oracle.

Perspective

The result targets self-distillation-style policy optimization with a privileged teacher, in two training settings: mathematical reasoning with Qwen3-4B on DAPO-17K, and multimodal reasoning with Qwen3-VL-8B-Instruct on MMFineReason-123K. Methodologically, once step-level direction and magnitude are fixed, teacher evidence only handles within-step token allocation, which reduces sensitivity to preference changes such as teacher wording or how answers are presented; the paper reports that DCSD costs more per training step than OPSD, though generation remains the dominant cost so the end-to-end increase is limited. For a reader, this offers a reusable design principle: treat "which direction to update" and "how much credit to give" as separate problems rather than letting one teacher signal decide both.

The paper treats a condition-matched Qwen3-32B as an operational oracle, so TDR and MAE measure agreement with a stronger-model reference rather than with the true oracle; the authors also note that "certified" refers only to probe-selected steps satisfying Theorem 2's error-budget condition, not to all empirical overrides. In multimodal experiments only DCSD is trained while baseline results come from published RLSD settings, and DCSD underperforms RLSD on ZeroBench, which the authors attribute to errors from fine-grained visual perception and spatial reasoning that credit calibration alone cannot address. In addition, figures such as the roughly 6% direction-correction rate and the roughly 1.5-times magnitude reduction come from training-log statistics whose aggregation is defined in Appendix E.3; readers who only see the abstract and main tables would need the appendix to confirm those conventions. Longer-horizon and interactive reasoning are listed as future work and are not yet demonstrated.

Sources