Skip to main content
Back to timeline
arXivSource publication:

Tsinghua LeapLab replaces KL divergence with a binary directional reward, matching or beating on-policy distillation on math and code, and proposes C-MOPD for multi-teacher training

Synopsis

This work revisits whether reverse KL divergence is necessary for on-policy distillation (OPD), proposes BinaryOPD, which keeps only the update direction as a binary reward and matches or slightly exceeds OPD across multiple student-teacher pairs on math and code, further finds that only the direction of a small subset of high-disagreement tokens is critical, and builds on this to propose C-MOPD, in which every sample is supervised by all teachers and which consistently outperforms MOPD on math and code benchmarks.

AI-generated editorial illustration: Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?

Interpretation

The authors propose BinaryOPD: assign reward +1 to a token when the teacher probability is higher than the student probability and -1 when it is lower, keeping only the update direction and discarding KL magnitude, and optimize the student with a policy gradient objective. Classical knowledge distillation and OPD default to reverse KL divergence as the loss; this work reduces the loss to a sign-level directional signal and shows this simplification reproduces the OPD training mode. On math, five student-teacher pairs are used (e.g., Qwen3-4B-Non-Thinking vs. Qwen3-4B-Non-Thinking-RL-Math, BinaryOPD average 65.1 vs. OPD 64.1; Qwen3-1.7B-Base pair 22.2 vs. 21.2); on code, two pairs (55.2 vs. 55.4), with training dynamics matching in accuracy, entropy, and response length.

The authors find that only the direction of a small subset of tokens with large teacher-student disagreement is critical: keeping only high-disagreement tokens and ensuring their update direction is toward the teacher reproduces OPD, while reversing their direction makes training fail. Prior practice assumes matching the teacher across the full vocabulary distribution; this work localizes the critical ingredient to a small set of high-disagreement tokens and shows low-disagreement tokens can be pushed away from the teacher without severe harm. Symmetric positive and negative reward experiments partition tokens into Groups A, B, and C by log-ratio threshold: keeping only Group A (a small fraction of all tokens) trains, while adding Group C makes training fail; in the negative experiment keeping only Group C trains and adding Group A fails; Group B, the vast majority of tokens, is not critical.

The authors propose Consensus Multi-Teacher On-Policy Distillation (C-MOPD): every sample is supervised by all teachers; when teachers agree the student is pulled toward all of them, and when they conflict it falls back to the domain-specific teacher and adopts its direction only if the conflict is not strong, otherwise assigning a zero reward. MOPD routes each sample to a single domain-specific teacher, so cross-domain updates may degrade other domains' capabilities; C-MOPD replaces routing with consensus supervision across all teachers to avoid such capability conflicts. With a Qwen3-4B-Non-Thinking student and math and code teachers on a mixed training set, C-MOPD outperforms MOPD on almost all math and code benchmarks, e.g., LiveCodeBench v6 32.9 vs. 27.9; an overly conservative threshold makes the zero-reward ratio rise monotonically beyond a certain level and performance only matches MOPD.

The authors explain why BinaryOPD works through state redundancy: states visited during OPD training are highly redundant, and the same or nearby states are revisited, so repeated directional feedback can determine the required correction iteratively without encoding magnitude in any single update. This explanation elevates the replaceability of KL magnitude from an empirical observation to an account of the closed-loop feedback mechanism in OPD, noting that state exploration saturates quickly. Recording 30 training steps and 61,440 states with a Qwen3-4B student and Qwen3-4B-Non-Thinking-RL-Math teacher on DeepMath, cross-prompt nearest-neighbor cosine similarity has mean 0.893 and median 0.901; the 2,048 states from step 1 already cover a substantial fraction of later states, and expanding the library about 29-fold yields only modest coverage gains.

Perspective

The results target researchers and engineering teams doing LLM post-training with on-policy distillation, in settings where a student model receives token-level teacher supervision on its own rollouts, such as mathematical reasoning and code generation; BinaryOPD and C-MOPD can directly replace the reverse KL loss or MOPD's routing mechanism, and code is released at LeapLabTHU/KL-Free-OPD. The C-MOPD experiments use a Qwen3-4B-Non-Thinking student with math and code domain teachers, and a training set mixing DeepMath, CodeContests, TACO, APPS, and Codeforces.

Identifying high-disagreement tokens depends on a threshold; the text reports supplementary experiments with different thresholds and another student-teacher pair, but how the threshold shifts across models and tasks remains an open question; the state-redundancy analysis relies on a fixed teacher representation space and a single student-teacher combination, so whether its conclusions hold under other representations or scales deserves continued observation; performance is largely unaffected when the verifier signal contradicts the update direction, and the practical trade-offs when combining with RLVR still need testing in more settings; in addition, some numerical values are missing in the parsed text, so specific thresholds and ratios require consulting the original tables and appendices.

Sources