Selecting distillation positions by entropy shift lifts LLM judges 2–9 points over Dr. GRPO on subjective tasks
Synopsis
This work studies training LLM judges from natural language feedback, proposes using the per-position entropy shift between teacher and student to distinguish context sharpening from context spreading, and masks high-entropy-shift positions so that self-distilled judges outperform outcome-supervised RL (Dr. GRPO) by 2–9 percentage points on the evaluated subjective subcategories while remaining competitive on objective ones.
Interpretation
The paper defines the per-position entropy shift as the reduction in the teacher's next-token entropy (teacher additionally conditioned on language feedback) relative to the student's, and uses it to identify two regimes: context sharpening at large positive shifts, where the teacher concentrates probability on a particular feedback-aligned criterion expression, and context spreading at large negative shifts, where the teacher distributes probability across multiple feedback-aligned alternatives. Prior outcome-supervised RL (GRPO, Dr. GRPO, DAPO) credits every token in a rollout with a single scalar determined only by final-verdict correctness, giving no separate credit at criterion-choice tokens and leaving the language feedback that accompanies preference labels unused; self-distillation offers per-position supervision but treats all positions as equally valid targets. Across 102 HelpSteer3-Preference validation rollouts and 104,046 response positions, per-position reverse KL shows a U-shaped relationship with entropy shift, with each tail (bottom and top 30%) carrying roughly 20–28 times the per-position KL of the middle 40%; criterion-selection spans are 16.7% of tokens but 44.4% of reverse-KL mass, and criterion-name positions are 0.4% of tokens but 11.1% of KL mass, with 94.8% falling in the two extreme 30% tails.
The authors propose position masking by entropy shift: rank positions within each generation by entropy shift, mask the top 70% with the largest values, and retain the bottom 30%, thereby favoring context-spreading positions. This turns the two-regime interpretation into a simple per-generation mask that needs no extra annotation, and setting the mask fraction to zero recovers naive self-distillation. A mask-fraction sweep on Qwen3-4B-Instruct and Qwen3-30B-A3B-Instruct raises total average from 80.17 (naive SD) to 82.50 at 30B and from 77.70 to 79.04 at 4B; direction-flip, both-tails, random, student-entropy, and teacher-entropy selectors at matched mask fraction all score lower.
On subjective subcategories both self-distillation variants gain 2–9 percentage points over Dr. GRPO, while Dr. GRPO remains competitive on objective subcategories. This complements the picture that outcome-supervised RL works well when the decisive criterion is clear-cut: when the verdict hinges on which criteria are invoked and how they are weighted, dense per-position supervision from language feedback has the advantage. At 30B, Chat 74.07 vs 76.74/80.53, Factuality 71.79 vs 78.53/79.16, Focus 80.00 vs 86.26/85.66; at 4B, Chat 68.91 vs 75.71/77.95, Factuality 62.74 vs 68.42/70.74, Focus 75.76 vs 80.20/77.78; on the objective side, 30B RM-Bench Math is 95.59 (Dr. GRPO) vs 94.45/95.63 and RewardBench v2 Math is 87.43 vs 81.42/84.15.
Masking high-entropy-shift positions is associated with broader criterion diversity and yields higher win rates in downstream pairwise selection. The paper extends the mechanism interpretation to inference-time behavior and downstream use: SD+mask invokes more distinct criterion clusters than naive SD at every evaluated clustering threshold and selects responses preferred under downstream evaluation in a multi-round best-of-eight tournament. Criterion diversity is measured over 300 samples each from RM-Bench, RewardBench v2, and HelpSteer3-Preference, with the difference growing from 4 clusters at one threshold to 26 at another; creative-writing win rates are 0.556 for Dr. GRPO, 0.612 for naive SD, and 0.632 for SD+mask.
Perspective
The method targets instruction-tuned models trained as judges, in settings where each preference example carries a one- or two-sentence rationale that names the decisive criterion; the authors also show masking still helps when generated rubrics serve as feedback, so the benefit is not tied to one feedback format. For teams that want dense supervision without extra annotation by reusing rationales already present in preference data, this offers a concrete path; for teams needing a general-purpose judge across diverse prompts, criterion diversity is a directly relevant property.
The authors state two scope limits: the method depends on feedback quality, and entropy-shift masking does not guarantee useful supervision when feedback fails to identify a decisive criterion or gives only generic guidance; the method also does not explicitly address known self-distillation issues such as hallucination and training instability in long-chain-of-thought reasoning models, which the authors leave to future work. In addition, the downstream validation uses a multi-round best-of-eight tournament rather than full GRPO policy updates, so the claim that it is a more reliable reward selector still needs confirmation inside a complete RLHF run.
