Cross-tokenizer on-policy distillation: strict alignment already covers 85.57–96.98% of student tokens, a student-selected top-16 subset retains at least 96% of the gain, while adding span supervision on mismatch groups lowers accuracy
Synopsis
Across three heterogeneous teacher–student pairs on mathematical reasoning and code generation, this study finds that strict alignment already covers 85.57–96.98% of student tokens despite static vocabulary Jaccard overlap of only 39.49–64.87%, that the shared vocabulary retains nearly all predictive probability mass at strict positions, that restricting reverse KL to a student-selected top-16 subset of the shared vocabulary preserves at least 96% of the full-shared-vocabulary improvement, and that adding span log-probability MSE on mismatch groups achieves complete supervision coverage yet lowers accuracy in all 18 positive-weight settings.
Interpretation
Strict alignment already covers most tokens on student-generated trajectories, so static vocabulary overlap understates the comparable supervision available. Prior cross-tokenizer distillation largely pursued broader alignment coverage; this work characterizes the comparable supervision space through dynamic coverage rather than static vocabulary overlap, reporting complete runs and three disjoint windows for three model pairs. Three teacher–student pairs (Qwen→Llama, Granite→Phi, Granite→Qwen), 100 iterations with 512 rollouts each, 51,200 responses per run; strict student-token coverage 85.57–96.98% and teacher-token coverage 82.82–97.26%, with window-to-window variation of at most 1.43 and 3.75 percentage points respectively.
The shared vocabulary carries nearly all predictive probability mass at strict positions, and a student-selected top-16 subset retains most of that mass for both models. Separates vocabulary mismatch from probability-mass distribution: static mismatch does not determine how much mass the shared vocabulary retains, and a subset selected using only the student distribution also covers most of the teacher's mass. Measured on the complete, unfiltered 20k prompts with pre-distillation student rollouts; the shared vocabulary retains 99.69–99.90% of teacher mass and 98.99–99.81% of student mass on average; top-16 retains at least 93.54% of teacher and 94.55% of student mass; expanding from 16 to 128 entries adds at most 3.44 points for the teacher and 2.70 points for the student.
Restricting reverse KL to a student-selected top-16 subset retains most of the learning benefit of full shared-vocabulary distillation, and all three strict variants exceed the evaluated cross-tokenizer baselines on math and code averages. Shows that concentrated supervision remains sufficient under heterogeneous tokenizers, without widening the supervision set for coverage; top-16 leads the strongest alternative by 0.51–1.05 percentage points in full average. Evaluated on MATH500, GSM8K, AIME-2024/2025/2026, AMC23, Minerva-Math and HumanEval, MBPP, LiveCodeBench with mean@32 for math and mean@8 for code; top-16 retains at least 96% of the Base-to-strict-full full-average improvement; in Appendix D with a 235B teacher to an 8B student, top-16 attains the highest Math, Code, and Full averages, and in Appendix E on ALFWorld it attains the highest success rate on both seen and unseen splits.
Applying span log-probability MSE to mismatch groups achieves 100% structural supervision coverage yet consistently lowers downstream accuracy, with gradient diagnostics offering an interpretation. Treats coverage and supervision reliability as separately measurable quantities: the value of added supervision depends on its interaction with strict distribution matching rather than on how many positions it covers. All 18 positive-weight settings score 0.27–1.20 percentage points below their strict baselines; at checkpoints 0, 20, 60, and 100 from strict-only training, span gradients show weak or negative directional agreement with strict gradients, consistently below a split-half reference formed from strict positions, and the span-to-strict gradient norm ratio increases across checkpoints.
Perspective
The result targets practitioners performing on-policy distillation with heterogeneous tokenizers: when teacher and student vocabularies differ but many strictly aligned positions exist on student trajectories, reverse KL at strict positions with a support compressed to the student-selected top-16 shared-vocabulary subset is the recommended configuration. The applicable setting is the mathematical reasoning and code generation studied here, plus the larger-teacher and ALFWorld embodied tasks in the appendices; training uses 100 on-policy iterations, one response per prompt, reverse KL, temperature 1.0, with evaluation as mean@32 for math and mean@8 for code.
The weak directional agreement between span and strict gradients, together with the rising relative span gradient norm, is the authors' diagnostic interpretation of the accuracy drop rather than a causal proof; whether the same explanation holds in the larger-teacher and ALFWorld settings is not accompanied by matching gradient analysis in the text. Probability-mass measurements use one pre-distillation student rollout per prompt and exclude empty responses, responses whose teacher and student tokenizations yield different byte streams, responses without scorable strict positions, and over-length inputs, so the retained-mass figures are averages over scorable strict positions. There is no consistent gain between top-16 and top-128, leaving the optimal subset size open across model pairs and tasks. The study centers on mathematics and code, so behavior in other domains and longer-generation settings remains to be examined.
