Skip to main content
Back to timeline
arXivSource publication:

Peking University-led team distills coding ability from one-word answers: a Qwen2.5-1.5B student beats an exact nuisance-matched control by 5.34 points on HumanEval+

Synopsis

The work introduces Active Taskless Distillation (ATD), which uses a public ancestor model to select unrelated prompts on which it is nearly indifferent between two ordinary words, has a privately post-trained teacher return a single word per prompt, and trains a same-ancestor student only on those prompt-word pairs; in the primary Qwen2.5-1.5B coding experiment, 5,664 single-word responses yield a 5.34 percentage-point gain on HumanEval+ over an exact nuisance-matched control, with positive mean gains also observed in scientific knowledge, commonsense reasoning, and reading comprehension across additional model generations, sizes, and families.

AI-generated editorial illustration: Post-Training Leaves Behavioral Shadows on Unrelated Decisions

Interpretation

ATD treats the behavioral shadow of a post-training update on unrelated inputs as an observable object: it uses the public ancestor to select prompts where two ordinary words are nearly equally likely, has the teacher return one word per prompt, and trains the student solely on the resulting prompt-word pairs. Prior subliminal-learning work largely transmitted traits or preferences and relied on extensive teacher outputs; this work compresses each observation to a single word and reports capability-level transfer. The primary experiment uses Qwen2.5-1.5B-Instruct as both public ancestor and student initialization, with a frozen rank-16 code-DPO LoRA teacher; 5,664 single-word responses form the entire training signal, and a frozen audit finds no code, mathematics, benchmark, or task terms.

In the primary coding experiment the student matches the teacher's aggregate coding score and exceeds every control: by 5.34 percentage points over the exact nuisance-matched permutation, 5.03 over the global teacher-label shuffle, and 4.57 over the teacher-free public-label arm. The exact control holds the completion multiset and per-bin flip counts fixed and only breaks the prompt-token correspondence, so the gain cannot be explained by unigram frequency or the teacher's label marginal. HumanEval+ is scored by greedy pass@1 through execution, which a student cannot fake by imitating carrier words; intervals are bootstraps of paired per-task signal-minus-control differences crossing training seed and task.

The effect persists when the acquisition is re-generated and the teacher re-trained, and across adaptation regimes: five independently constructed acquisitions, each with its own teacher and exact control, crossed with three training seeds, give an aggregate effect of +4.80 percentage points with 15/15 and 5/5 positive, and LoRA r16, r32, r64, and full-parameter SFT students all beat their exact nuisance-matched controls. The main result no longer varies only the student's optimization seed; it resamples the acquisition, the teacher, and the training seed together. The aggregate interval is a three-axis crossed bootstrap with a between-acquisition standard deviation of 0.9 percentage points; per-regime signal-minus-control gaps are +5.34, +6.55, +9.15, and +4.67 percentage points.

The transfer is source-specific and composable: across three sources and three target endpoints each shadow has its largest effect on its matched target with off-diagonal effects near zero or negative, a 50/50 mixture of math and code shadows recovers both sources, and the recovered direction tracks the teacher's update strength. The channel is not a generic 'train harder' direction; it carries source-relevant capability information, and multiple sources can superpose. Source specificity comes from the three-source, three-endpoint matrix; the mixture shows math and code alignment advantages of +0.046 and +0.062, each positive in 3/3 seeds; graded strength recovery has Spearman 1.0 on both the code and SciQ trajectories.

Perspective

The result applies in a known-ancestor setting: the teacher is obtained by privately post-training the public model, the student is initialized from the same ancestor, the teacher is accessible only through a black-box interface returning one greedy token per query, and the query set must be fixed before any teacher response is observed. It speaks to researchers studying the observability of post-training updates and implicit distillation, and to service operators who need to assess whether private adaptations leak information through unrelated behavior. Carriers are selected by the public ancestor without teacher information, and the 5,664 training rows of the primary coding acquisition are audited to contain no code, mathematics, benchmark, or task terms, so the instrument can test capability transfer without target-task examples, teacher logits, or teacher parameters.

Transfer is not reliable in every tested setting: Appendix H reports no transfer under an incompatible ancestor, and the recovered effect is near zero on the MBPP+ endpoint where the teacher's own gain is small, which the authors read as a low-bandwidth channel. The authors also note that the current approach relies on actively constructed carriers and a known public ancestor, leaving open whether the selected observations reveal too little capability-relevant information or whether the learning procedure fails to make sufficient use of it. The multi-token observation gives a large effect on the arithmetic endpoint but small, sign-consistent effects elsewhere with intervals that reach close to zero at that scale. In addition, this document is the full text, but some tables and figures are rendered as text and a few numeric spans are incomplete in the loaded markdown, so readers needing exact reproduction should consult the original tables.

Sources