Skip to main content
Back to timeline
arXivSource publication:

LinkedIn and university collaborators find that 7B-to-72B students trained on rejects from smaller frozen models do better on code and math while using less compute

Synopsis

Holding prompts, chosen responses, SeqKD initialization, DPO objective, and training budget fixed while varying only the reject source, the study finds that across Qwen2.5 students from 7B to 72B, rejects generated by smaller frozen Base models with less inference compute train stronger students than the students' own rejects on code generation and mathematical reasoning, and it derives a finite-horizon utility bound in a linearized feature model that motivates three interventions: source mixing, prompt reassignment with lexical permutation, and reference-likelihood reselection.

AI-generated editorial illustration: Smaller Models, Better Rejects: Preference Distillation Scaling

Interpretation

The work establishes a cross-scale preference-distillation phenomenon: at all four student scales of 7B, 14B, 32B, and 72B, every smaller frozen Base source beats both vanilla Self and post-SeqKD Self reject constructions on code avg@4, and on math reasoning most of the 14 Qwen configurations with a smaller Base source exceed Continued-SFT. Prior preference distillation treats the teacher response as preferred and the student's own response as rejected, implicitly assuming self-generated failures are the most informative negatives and that rejects must come from a model at least as large as the student; this work treats the reject-generating distribution as an independent design variable and changes only the reject source under a shared protocol. The code line uses 75,752 SeqKD prompts and a disjoint 24,248-prompt preference split from KoDCode, with every reject verified incorrect by failing at least one unit test, evaluated on 5,000 held-out problems plus BigCodeBench, HumanEval, and MBPP; the math line uses 33,524 SeqKD prompts and 12,288 preference prompts from OpenR1-Math, evaluated by avg@4 on MATH-500 and AIME.

The work casts reject source selection as inverse data design, derives a finite-horizon utility bound in a linearized feature model of DPO, characterizes each source by transfer mass and adverse share, and states a favorable-region condition. Existing work largely selects or weights observed pairs or couples data generation with training, whereas this analysis characterizes which reject distribution improves a given student's task utility and derives three testable construction hypotheses from it. The derivation uses network score gradients at the SeqKD parameters as fixed response features and linearizes around the shared DPO initialization; the proposition gives an explicit correction term under regularity conditions such as bounded features, with full construction and proofs in the appendix.

Three interventions are supported: randomly mixing smaller-Base and vanilla-Self rejects raises code avg@4 as the smaller-Base share grows; reassigning real Base rejects to other prompts, and even keeping only shuffled code tokens, still beats length-matched gibberish; and reselecting within a fixed candidate bank by lower reference likelihood beats higher-likelihood selection for every source. These interventions turn the theory's transfer mass, task structure, and reference coupling into operational construction rules, and show that part of the gain comes from task structure itself rather than from errors tied to each specific prompt. In the mixing experiment the 14B student's avg@4 rises from 59.730 to 63.280 as the Llama-1B share grows and from 59.575 to 62.630 for Qwen2.5-3B; reassignment and lexical permutation beat gibberish at every student scale; the reselection experiment compares Repelled, Native, and Attracted selection on a shared set of 20,136 prompts, with the 1.5B source reaching 63.135 under Repelled versus 61.880 under Attracted.

The work identifies two observable properties of useful rejects: preserving task-relevant structure while limiting coupling to the DPO reference policy, both of which smaller frozen models provide at lower generation cost. Natural-source analysis shows higher reference likelihood associated with lower downstream avg@4, with smaller Base sources systematically on the low-coupling high-utility side and both Self constructions on the high-coupling low-utility side; reselection further supports treating reference coupling as an operational property of the reject distribution rather than only a proxy for model scale. The cost proxy puts smaller Base sources at roughly 0.9% to 50.1% of same-scale vanilla Self total cost, a 49.9% to 99.1% reduction, with a matched H100 runtime sweep following the same direction; optimization diagnostics including reject fields, reward margins, likelihood trajectories, and RC-DPO do not recover the smaller-Base-over-Self ordering.

Perspective

The result applies to two-stage preference distillation combining SeqKD and DPO in settings with executable or verifiable answers, such as code generation and mathematical reasoning; in such settings practitioners can first generate rejects with a smaller frozen Base model, then select low-likelihood candidates under the reference, and mix in a growing smaller-Base share. Cross-family Llama sources beat Self in all nine code configurations and in eight of nine MATH-500 configurations, suggesting the practice is not limited to the student's own model family. The theory is a finite-horizon bound in a linearized feature model, valid under bounded features and sufficiently small step sizes.

Why reject utility survives token shuffling is attributed to code-domain lexical structure, but what exactly that structure contains remains an open question. The negative association between reference likelihood and downstream utility holds for natural sources and reselection supports its operational use, yet the causal chain relies on the linearized model's approximation. Optimization diagnostics fail to recover the source ordering, indicating current diagnostics are not yet sufficient to predict reject-source quality before training. The text was read in full, but some table values are missing from the evidence bundle, so exact endpoints still require the original appendix.

Sources