Skip to main content
Back to timeline
arXivSource publication:

BiToK-SD learns Top-K token selection via bilevel optimization, achieving the best average scores across three scales in mathematical-reasoning self-distillation

Synopsis

The work proposes BiToK-SD, which casts Top-K token selection in on-policy self-distillation as a bilevel optimization problem: the lower level turns discrete Top-K selection into differentiable soft weights via a threshold relaxation that adapts as the student policy evolves, while the upper level performs knowledge distillation only on the selected positions; trained for 200 steps on roughly 30K problem-solution pairs from the mathematical-reasoning subset of OpenThoughts with Qwen3-1.7B/4B/8B, it attains the highest average scores among all compared methods on AIME24, AIME25, and HMMT25 under Avg@64, exceeding the strongest baseline OPSD by 0.63, 0.58, and 0.21 average points respectively.

Source-provided article image: Learning What to Distill: Bilevel Top-K Token Selection for Self-Distillation in Large Language Models
Figure 1 ·

Figure 1: Exploration dynamics during training. (a) Student entropy over training steps. (b) Student–teacher entropy gap over all tokens. (c) Student–teacher entropy gap on selected tokens. BiToK-SD maintains higher student entropy and larger student–teacher entropy gaps than the baselines, suggesting stronger policy exploration.

arXiv

Interpretation

It formulates token selection in self-distillation as bilevel optimization: the lower level approximates Top-K selection with a differentiable threshold relaxation, and the upper level minimizes a weighted self-distillation loss on the selected token positions, thereby learning both where to distill and how strongly. Earlier selective distillation either distills all tokens uniformly or uses fixed heuristic scores such as entropy and teacher-student divergence for binary Top-K selection, assigning identical weight to selected positions; BiToK-SD lets the selection threshold update online with the student policy and yields continuous rather than binary token weights. The method is fully specified in Section 3, including the utility score, relaxed threshold objective, soft mask, and weighted loss; ablations show mask density stays near the target sparsity and the threshold keeps changing during training (on 1.7B from 0.061 to 0.063, range 0.044-0.105; on 4B from 0.030 to 0.102; on 8B from 0.069 to 0.159).

Across three model scales on mathematical reasoning benchmarks, BiToK-SD achieves the highest average performance among all compared methods and is best in seven of the nine scale-benchmark cells. Relative to the strongest baseline OPSD, average scores improve by 0.63, 0.58, and 0.21 points at 1.7B, 4B, and 8B; relative to TIP, which uses the same Soft-OR score but a hard Top-K mask and reverse KL, it leads in all nine cells. Evaluation uses Avg@64 (mean accuracy over 64 independently sampled solutions per problem) on AIME24, AIME25, and HMMT25 under thinking mode; all methods train for 200 steps and are evaluated at the final checkpoint, deliberately avoiding test-set-based checkpoint selection, with all methods initialized from the same checkpoint and trained on the same data.

The learned selector keeps the prescribed sparsity budget while assigning larger weights to reflection-associated token positions and preserving stronger policy exploration. OPSD weights all valid positions uniformly and TIP applies a binary Top-K mask, so in both baselines the relative weighting between reflection tokens and other selected tokens is uniform by construction; BiToK-SD's continuous weights increase with the utility score relative to an adaptive threshold. Using ten epistemic markers (wait, hmm, perhaps, maybe, actually, alternatively, seems, might, likely, check) as lexical indicators of self-reflective behavior, BiToK-SD generates more markers per completion than OPSD and TIP, with the gap appearing around step 60 and widening; reflection-associated positions receive larger mean soft-mask weights than other valid positions; student entropy and the student-teacher entropy gap are larger, and on AIME24 it generates 3.7%, 1.6%, and 4.8% fewer tokens than OPSD with lower repeated 4-gram ratios and truncation rates.

The bilevel selection mechanism transfers to off-policy distillation: it still outperforms full-position distillation and student-entropy-based selection when only truncated teacher distributions are available. Because storing the full teacher distribution is prohibitively expensive in the off-policy setting, only top-64 probabilities are cached, so the full-distribution KL terms required by Soft-OR cannot be faithfully computed; the student-entropy score, which depends only on student predictions, is used instead, with the selection ratio fixed at 10%. Following the protocol of Tavor et al., Qwen3-8B is distilled into Qwen3-1.7B on an 80M-token subset of FineWeb for one epoch with roughly 2.4K optimizer updates; BiToK-SD outperforms Full-KD, SE-KD, and SE-KD3X on all three summary metrics: average accuracy over ARC-Easy, GSM8K, HellaSwag, PIQA, and LAMBADA, the strict instruction-following score on IFEval, and LAMBADA perplexity.

Perspective

The result targets self-distillation for compressing reasoning models in resource-constrained settings: it applies to Qwen3-1.7B/4B/8B trained for 200 steps on roughly 30K problem-solution pairs from the mathematical-reasoning subset of OpenThoughts and evaluated with Avg@64 on AIME24, AIME25, and HMMT25 under thinking mode, for practitioners seeking to improve small-model mathematical reasoning without an external large teacher. The method is not tied to a specific utility score and can use student entropy, teacher entropy, forward KL, or reverse KL; the off-policy experiment further shows that when only cached top-64 teacher probabilities are available and full-distribution KL cannot be computed, switching to the student-entropy score still yields gains in general-corpus distillation.

The choice of utility score and selection ratio still affects the gain over uniform weighting: in ablations Soft-OR and forward KL beat uniform weighting, whereas entropy-based scores and reverse KL do not, with differences across utility scores within about two Avg@64 points; the Top-K ratio follows an inverted-U trend peaking at 50%, indicating the benefit depends on intermediate sparsity. TIP and BiToK-SD differ in both the selection rule and the divergence, and the text notes the gap cannot be attributed to either factor alone. In the off-policy setting only truncated teacher distributions are available, so Soft-OR cannot be faithfully computed and that conclusion rests on the student-entropy score. Reflection behavior is measured indirectly through ten epistemic marker words, and the main experiments are limited to mathematical reasoning benchmarks and the Qwen3 family, leaving applicability to other reasoning domains and model families an open question.

Sources