Separating Teacher Self-Deviation from the Distillation Signal: Calibrated On-Policy Distillation
Synopsis
The work argues that the token-level teacher–student discrepancy used by on-policy distillation (OPD) mixes in context-induced teacher-side variation, termed Teacher Self-Deviation (TSD), and introduces Cal-OPD, which estimates the teacher's self-deviation region via positive and negative privileged interventions and keeps only the residual beyond that region as the optimization signal, consistently outperforming standard OPD and its variants on mathematical reasoning benchmarks while retaining only about 52–65% of the original discrepancy.
Interpretation
It identifies and characterizes Teacher Self-Deviation (TSD): with the problem, student rollout, and evaluated token held fixed, changing only the teacher-side context still substantially shifts the teacher's log-likelihood for the same token. Prior work treated teacher likelihood as a uniformly reliable pointwise reference, or discussed supervision reliability through uncertainty, position, outcomes, or trajectory quality; here the variation of teacher likelihood under controlled contextual interventions is explicitly modeled as an estimable deviation region. With Qwen3-1.7B as student and Qwen3-8B as teacher, rollouts were generated on 6,528 questions sampled from DAPO-17k, and roughly 60 million response tokens were rescored under a baseline plus eight contextual interventions, with reproduction across thresholds and teacher scales (Qwen3-4B-Thinking-2507).
TSD emerges substantially without task knowledge, is largely insensitive to intervention semantics and correctness, and concentrates on surface-form tokens rather than mathematical content. This indicates that teacher likelihood shifts cannot be reliably read as task knowledge, and that richer privileged context mainly broadens the already-affected token positions rather than creating an entirely new deviation pattern. Under task-agnostic instructions the union prevalence of significant TSD is comparable to evaluative feedback and answer-level privilege; positive and negative variants show high positional overlap, directional agreement, and shared deviation ratios; the highest-TSD forms are maybe, however, therefore and similar markers, while the lowest are digits and mathematical symbols, a nearly order-of-magnitude separation.
It proposes Cal-OPD, which estimates the teacher's self-deviation region from positive and negative privileged interventions, zeroes the portion of the teacher–student discrepancy covered by that region, and keeps only the residual beyond its boundary as the token-level advantage. Privileged information is used to calibrate the teacher reference rather than being distilled directly into the student, in contrast to Privileged-OPD, which distills the privileged teacher's output directly. An explicit decomposition and a relaxation factor ε for the region estimate are given, and evaluation uses Avg@16 on two teacher–student scale configurations (4B→1.7B and 30B-A3B→4B) across six mathematical reasoning benchmarks.
On mathematical reasoning benchmarks, Cal-OPD retains only about 52–65% of the original teacher–student discrepancy yet achieves the highest average performance in both configurations and mitigates the response-length expansion seen in standard and privileged OPD. Standard OPD improves the 4B→1.7B student only modestly and even degrades the 30B-A3B→4B student, while Privileged-OPD has the lowest average among distillation methods in both configurations; retaining less signal works better, suggesting that learning more discrepancy is not the key. Cal-OPD has the highest average in both configurations and the best result on 7 of 12 benchmark–configuration pairs; ablations show evaluative-feedback interventions perform best while solution-level privilege over-filters and lowers the average; retention-matched controls (Advantage-Sync, Token-Sync, TSD-Filter) do not reproduce its gains.
Perspective
The result targets research and engineering settings that use on-policy distillation for reasoning post-training: teacher and student from the same model family, mathematical reasoning tasks, Avg@16 accuracy as the metric, and calibration that relies on constructible positive and negative privileged interventions such as evaluative feedback, answer-level, or solution-level privilege. It supports using privileged information to estimate the teacher's self-deviation region and distilling only the residual, rather than teaching the privileged teacher's output directly; when the teacher–student size gap is moderate, it can also shorten total training time by avoiding response-length expansion.
Estimating the calibration region depends on the chosen positive and negative interventions: evaluative feedback performs best, while solution-level privilege over-filters and causes signal collapse, so intervention choice remains a design dimension to explore. The optimal relaxation factor ε (5 in the text) and its relation to the retained ratio, as well as stability across teacher scales and tasks, remain open questions. In addition, although this is a full-text parse, some figures and appendix details are conveyed in prose, so reproducing specific numbers and curves would still require checking the original.
