Post-training capabilities leak through unrelated text: one teacher word per prompt yields a 5.34 pp gain on HumanEval+
Synopsis
The authors introduce Active Taskless Distillation (ATD), which selects prompts where the teacher and student's shared public ancestor is nearly indifferent between two ordinary words, and lets a student initialized from that ancestor learn solely from the resulting prompt-word pairs, without target-task examples, teacher logits, or teacher parameters; in the primary coding experiment with Qwen2.5-1.5B, 5,664 samples yield a 5.34 pp gain on HumanEval+ over an exact nuisance-matched control, with transfer also shown in scientific knowledge, commonsense reasoning, and reading comprehension across additional model generations, sizes, and families, and functional analyses indicating the learned signal is composable and tracks the teacher's update strength.
Figure 1 : Reading a post-training update through target-unrelated decisions with ATD . Using only the public base M 0 M_{0} , ATD samples ordinary prompts on which M 0 M_{0} is nearly indifferent between two words (near-ties), and asks the private teacher M T = M 0 + δ M_{T}\!=\!M_{0}\!+\!\delta for one word at each; where the update tips a near-tie, the returned word is a one-bit observation of the teacher’s behavioral shadow. A student initialized from the same M 0 M_{0} then learns only from these prompt–word pairs, and a matched control keeps the same words but breaks the pairing. Evaluated on held-out target tasks it never saw in training, the student trained on the true pairs outperforms the control, and different shadows lift different capabilities.
arXivInterpretation
Language models can transfer capabilities through task-unrelated text, so post-training leaves behavioral shadows on unrelated decisions. Prior work on subliminal learning focused largely on traits or preferences and relied on extensive teacher outputs; this work extends the phenomenon to capability transfer and compresses supervision to a single word per prompt. The abstract reports a primary coding experiment with Qwen2.5-1.5B in which 5,664 samples produce a 5.34 pp gain on HumanEval+, with an exact nuisance-matched control.
ATD constructs a training signal of prompt-word pairs by selecting prompts where the teacher and student's shared public ancestor is nearly indifferent between two ordinary words. The student needs no target-task examples, teacher logits, or teacher parameters, learning only from these prompt-word pairs. The method description states the student is initialized from the shared ancestor and learns only from prompt-word pairs; the primary experiment provides a quantitative comparison against a matched control.
The transfer is not limited to coding and also appears in scientific knowledge, commonsense reasoning, and reading comprehension, across additional model generations, sizes, and families. It broadens the observation of capability transfer from a single model and task to multiple task types and model configurations. The abstract states experiments were run on additional model generations, sizes, and families, but does not give per-task numbers in the abstract.
Functional analyses show the learned signal is composable and that its strength tracks the teacher's update strength. It offers an initial characterization of the internal structure of the behavioral shadow, beyond whether transfer occurs. The abstract summarizes this as functional analyses without listing specific metrics or experiment sizes.
Perspective
This work is aimed at researchers and engineers studying how post-training shapes model behavior, and applies to settings where a model is initialized from a public ancestor and one wants to transfer capabilities without target-task examples. It suggests that very lightweight prompt-word signals can probe the behavioral traces left by teacher updates on tasks such as coding, scientific knowledge, commonsense reasoning, and reading comprehension, and provides a starting point for studying transfer across model generations, sizes, and families.
The abstract does not give specific gains for scientific knowledge, commonsense reasoning, and reading comprehension, nor does it quantify what 'composable' means or how strength tracks the teacher's update strength in the functional analyses. The mapping between the 5,664 samples and the 5.34 pp gain, the exact construction of the matched control, and the scale of experiments across model generations, sizes, and families all need confirmation in the full text. In addition, the abstract text contains several incomplete passages, so details should be checked against the original.
