1% of Tokens Can Match Full Distillation: IER Re-ranks Sparse Supervision by Gradient-Estimation Reliability
Synopsis
The work reframes token selection in on-policy distillation (OPD) as a gradient-estimation reliability problem at a fixed prefix, decomposes the one-sample reverse-KL gradient into signal and sampling noise in information geometry, proposes the information-efficiency ratio (IER) as a signal-to-noise ratio under the optimal scalar baseline, and approximates it on a top-K candidate set for ranking; on mathematical and medical reasoning, IER alone approaches or exceeds full OPD at token budgets of 0.1%–1%, and combined with existing usefulness scores (IER-OR, IER-AND) it matches or exceeds full OPD in multiple settings.
Interpretation
The paper identifies an overlooked axis in sparse OPD: even when teacher supervision at a token is useful, estimating the reverse-KL gradient from a single sampled next token can yield a noisy update, so token selection should account for both usefulness and gradient-estimation reliability. Prior OPD token selectors (Prefix, Entropy, TIP, TA-OPD, CA-SoftOR and others) were designed around usefulness, importance, teachability, or position; the authors state these works "mostly emphasize different aspects of token usefulness and ignore the estimation reliability," and this paper adds reliability as a complementary second axis. The claim is supported by a derivation: at a fixed prefix, an action-independent scalar baseline is introduced, the gradient signal and mean-squared estimation variance are decomposed, and the variance-minimizing optimal baseline is derived; the authors also note that high IER does not indicate high usefulness of the supervision.
The authors propose the information-efficiency ratio (IER): under the optimal scalar baseline, the ratio of expected gradient signal to minimized sampling error, whose reciprocal corresponds to the relative mean-squared error of single-sample token gradient estimation. Unlike vOPD and related variance-reduction work on single-sample OPD gradients, this paper analyzes the one-sample reverse-KL gradient from information geometry (natural gradients under the Fisher metric) and provides a signal-noise decomposition with a closed-form optimal baseline. The result is stated as Theorem 1 with a proof in the appendix; Corollary 1 links the reciprocal of IER to the relative MSE of an average over independent samples and gives the sampling complexity needed for a target relative MSE.
To cut computation, the authors approximate IER on a candidate set built from student and teacher top-K logits plus the sampled token, producing a ranking proxy that can be combined with existing usefulness scores via soft OR (IER-OR) and soft AND (IER-AND), while retaining the sampled reverse-KL training objective. This turns a full-vocabulary statistic into a per-prefix ranking score that plugs into existing sparse-OPD pipelines rather than replacing the training objective. Implementation details are given in the appendix: K=10 for top-K extraction, small values assigned to missing logits on the candidate set with renormalization, and clipping of the log-likelihood ratio for numerical stability; each rollout batch contains four prompts with 16 student responses each, selector scores are normalized within the batch under one global budget, and at least one token is kept per response.
On mathematical reasoning (strong-to-weak JustRL-Nemotron-1.5B to OpenMath-Nemotron-1.5B, and big-to-small JustRL-Qwen3-4B to Qwen3-1.7B) and medical reasoning (ClinAlign-4B to Qwen3-4B), IER and its combinations match or exceed full OPD at very small token budgets of 0.1%–1%, and higher budgets do not always improve performance. The authors report that IER alone at a 0.1% budget reaches 58.9 vs. 59.9 for full OPD on AIME26 and 34.1 vs. 34.7 on HMMT26, and exceeds full OPD on three of four benchmarks for the other model pair; in medical reasoning at 0.1%, IER reaches a HealthBench overall score of 45.25 versus 45.77 for full OPD, while Prefix "provides essentially no improvement" at that budget. Evidence comes from experiments across multiple tasks, model pairs, budgets (0.1%, 1%, 10%, plus a sweep over 5%, 20%, 50%, 80%), and both thinking-on and thinking-off modes; the authors also state that gains depend on the selector and setting, that Entropy-based combinations sometimes reduce mathematical reasoning performance, and that both scores are approximations with no guarantee of better selection.
Perspective
The result targets on-policy distillation settings where a teacher scores each token of student-generated trajectories, covering verifiable mathematical reasoning and open-ended medical reasoning, both strong-to-weak distillation (same-size teacher post-trained on the task) and big-to-small distillation (larger, stronger teacher sharing the tokenizer), and both thinking-on and thinking-off modes. It provides a reliability ranking signal that can be inserted into existing selectors: IER can be used alone, or combined with usefulness scores such as Prefix, Entropy, TIP, TA-OPD, and CA-SoftOR through IER-OR (useful or reliable) and IER-AND (useful and reliable), while the training objective remains sampled reverse KL. The authors suggest that a truncation design based on IER, pruning a trajectory once enough tokens with reliable gradient estimates have accumulated, is a promising route to training efficiency, and leave it as future work.
The authors themselves note that how best to select and weight tokens remains open: gains vary across selectors and settings, Entropy-based combinations sometimes reduce mathematical reasoning performance, and both usefulness scores and IER are approximations with no guarantee of better token selection. IER measures relative estimation accuracy rather than gradient magnitude, and the authors state explicitly that high IER gives no guarantee about gradient magnitude, task alignment, or the benefit of the resulting parameter update, so IER is not a substitute for usefulness. In addition, their implementation still generates complete trajectories and adds computation for token scoring and selection, so sparse supervision does not imply proportional compute savings; the training-efficiency table shows IER and its combinations increase mean step time by about 2.2%–2.5% over full OPD and peak GPU memory by up to 2.17 GiB (0.92%). Medical reasoning is graded by gpt-oss-120B as the judge, which achieves a 0.6614 macro F1 on the HealthBench meta-evaluation, lower than GPT-4.1, a choice the authors attribute to cost. Finally, in the loaded markdown the numeric cells of the main mathematical reasoning table (Table 2) are empty, so per-selector scores on AIME25/26 and HMMT25/26 cannot be verified item by item; only the prose and the medical reasoning table can be used, and exact comparisons require the original tables.
