Sparse crosscoders show on-policy distillation adds no new student features but reweights features the student already shares with the teacher
Synopsis
Using sparse crosscoders and a proposed swap readout, the study reads how each student checkpoint uses every feature before and after on-policy distillation (OPD) across three OPD settings, finding that OPD neither creates the student's own features nor passes on the teacher's, but slightly reweights features the student already shares with the teacher (over 98% of frequently used features change firing rate by less than 20%), with the largest changes concentrated on decision tokens such as Wait, Hmm, and So; the SFT warm-up on the teacher's rollouts that commonly precedes OPD also adds no features, reweighting shared features partly along OPD's direction and partly beyond it, and imposing this reweighting on a directly distilled student's features without changing its weights brings its a
Interpretation
The authors propose the swap readout: it places a student checkpoint's activation in both student slots of the crosscoder while holding the teacher's input fixed, thereby reading that checkpoint's own feature activations even for checkpoints the crosscoder has never seen. Existing crosscoder analyses identify features specific to one model through their decoders but cannot show how a model's use of its features changes, since all models are encoded jointly into a single set of feature activations; the authors note the two students' activations differ by only a small fraction of their norm, making own-slot readouts unreliable. The paper gives a proposition and proof that the swap readout is invariant to a replacement matrix and equals the joint pre-activation whenever the two students coincide; across four three-model crosscoders it reconstructs students as well as the joint encoding and reconstructs checkpoints it was never trained on.
Across three OPD settings (JustRL, Skywork, R1-7B), OPD creates no features of the student's own and passes on none of the teacher's; it only slightly reweights the features the student already shares with the teacher. Prior analyses of OPD examine the student's outputs (token probabilities, accuracy, training signal); this work turns to the student's internal representations and offers a representation-level account: if OPD can only reweight shared features, it should improve sampling efficiency without expanding the capability boundary. In all three settings the OPD student dominates no feature and its MAS distribution coincides with the base student's; on M held-out tokens no feature is gained or lost, and 98.9%, 98.4%, and 99.9% of frequently used features change firing rate by less than 20% under JustRL, Skywork, and R1-7B; the student's MAS on teacher-dominated features is 0 both before and after OPD, and no feature's MAS changes by more than 0.01.
The features OPD changes most concentrate on decision tokens, the points at which a reasoning trace decides its next move, and OPD moves these features toward the teacher's usage. This localizes OPD's reweighting to specific positions along the trace and ties it to the steps where the teacher disagrees with the student most, which prior work had not characterized at the feature level. Decision-token features make up 40% of the most-changed features under JustRL and 30% under Skywork against a 4% baseline; at decision tokens the student replaces 1.5–2 times more active features than at an average token, and OPD's change to the next-token distribution is 2–3 times over-represented at those steps; the teacher–student KL divergence there is 2–3 times (JustRL), 3–4 times (Skywork), and 1.5–2 times (R1-7B) its average; under R1-7B OPD's average KL to the base student is only 0.01, a tenth of JustRL's.
The SFT warm-up on the teacher's rollouts that precedes OPD also adds no features; it reweights shared features, partly doing part of OPD's work in advance along OPD's direction and partly moving beyond it in ways that persist through OPD, and imposing this reweighting on a directly distilled student's features without changing its weights brings its accuracy close to the warmed-up student's. Unlike Shi et al. (2026), where SFT on solutions written by a separate, stronger model introduces new features, this warm-up stays within the task OPD then trains on and imitates the rollouts of the very teacher OPD distills from; the authors support the reweighting's role with a feature-level intervention rather than correlation alone. The warm-up moves features along OPD's direction with a slope of about 0.5, and 80% of the features OPD raises most move the same way after the warm-up; 60% of the warm-up's change lies outside OPD's direction (55% excluding one feature that fires on a garbled character), and this residual survives the subsequent OPD (Spearman 0.75, against 0.05 for a random rotation); in the intervention, the directly distilled student's avg@8 rises from 18.6 to 21.3 and the warmed-up student's falls from 21.9 to 19.2, while the shuffled-feature control gives 15.9 and 23.4.
Perspective
The work speaks to researchers studying OPD's internal mechanism and to teams designing post-training pipelines, and applies to settings where student and teacher share a feature dictionary and the task is mainly reasoning (math reasoning), for OPD and for the teacher-rollout SFT warm-up. The swap readout it introduces can compare how any checkpoint under a fixed crosscoder uses features, including checkpoints never seen in training, moving the question of what training changed from outputs to representations; the feature-level intervention also offers a way to test whether the reweighting carries the benefit without changing weights.
The conclusions rest on three OPD settings and one Qwen3 warm-up setting, and the authors state their warm-up conclusions do not apply to SFT in general; feature meanings rely on human reading of the tokens on which features fire most strongly, and the decision-token categories are the authors' grouping into reflection, hesitation, and progression. Whether mechanisms beyond reweighting operate, and how these findings hold at larger scales and on more tasks, remain open questions. In addition, this document is the full text, but some tables and figure captions appear as placeholders in the text, and a few numbers (such as some percentages and control values) are not fully rendered in the prose, so readers needing exact figures should consult the original tables and figures.
