Eight real prompts match 17k problems: data-free on-policy distillation hits full-data parity with 64 self-generated questions
Synopsis
Across two representative single-teacher on-policy distillation (OPD) settings, training on 8 real prompts yields performance comparable to training on 17k problems, and datasets differing substantially in measured difficulty and initial distillation gap yield similar outcomes; building on this, the authors propose a data-free OPD (DF-OPD) setting in which 64 self-generated questions obtained without seed examples match full-data OPD in both single-teacher settings, and in multi-teacher OPD 1k generated questions across mathematics, code, and instruction following match training on approximately 7k real post-training examples.
Figure 1: Reducing dependence on training questions. (a) OPD uses external questions. (b) DF-OPD uses teacher-generated questions. (c) Empty prompts work only in some tested settings; success is associated with student self-questioning and answering (implicit DF-OPD). OPD arrows denote token-level teacher supervision on student-generated trajectories.
arXivInterpretation
OPD depends far less on the number of training questions than expected: in two representative single-teacher OPD settings, training on 8 real prompts yields performance comparable to training on 17k problems. Prior work in this area focused largely on algorithmic advances, leaving the role of the training questions themselves uncharacterized; this work treats training-data quantity as an explicit variable with a comparable control. Evidence comes from controlled experiments in two single-teacher OPD settings comparing 8 real prompts against 17k problems; the text describes the outcome as 'performance comparable' without reporting specific metric values.
Differences in dataset difficulty and initial distillation gap have limited effect: datasets differing substantially in measured difficulty and initial distillation gap yield similar outcomes. It brings dataset difficulty and initial distillation gap, often assumed to be decisive factors, under direct test, indicating that OPD outcomes are not mainly determined by these dataset properties. Based on measurement and comparison of dataset difficulty and initial distillation gap; the text reports no specific measured values or statistical tests.
The authors offer two complementary explanations: repeated sampling could allow even a few prompts to expose substantial teacher supervision, while OPD transfers generalizable reasoning capabilities beyond dataset-specific knowledge. It provides candidate mechanism-level explanations for why little data suffices, linking the observation to sampling repetition and to capability transfer. The text frames this as 'Our analyses suggest', an analytical interpretation grounded in experimental observation rather than direct causal verification.
It proposes a data-free on-policy distillation (DF-OPD) setting: without seed examples, 64 self-generated questions match full-data OPD in both single-teacher settings, and in multi-teacher OPD, 1k generated questions across mathematics, code, and instruction following match training on approximately 7k real post-training examples. It pushes the reduction of data dependence to the point of requiring no external data at all, and shows the setting extends to multi-teacher, multi-task scenarios. Evidence consists of comparisons in both single-teacher and multi-teacher settings; the multi-teacher case spans mathematics, code, and instruction following, benchmarked against approximately 7k real post-training examples.
Perspective
The results target research and engineering settings that use on-policy distillation for post-training, particularly single-teacher and multi-teacher configurations; they indicate that on tasks such as mathematics, code, and instruction following, training questions can be sharply reduced or even self-generated by the system. For teams seeking to lower data-preparation cost or lacking access to high-quality external data, this offers a viable path of substituting self-generated questions for external data. The text also explores operating without explicit training questions at all, but this is effective only in limited cases, where the student unexpectedly generates and answers its own questions, reducing the process to an implicit form of DF-OPD.
The text describes several key comparisons as 'comparable' without specific metric values, benchmark details, or statistical significance, so the magnitude and stability of that parity still require judgment alongside the code and follow-up experiments. The two explanations (repeated sampling exposing teacher supervision, and OPD transferring generalizable reasoning capabilities) are analytical conjectures not yet directly verified. The case with no explicit training questions is effective only in limited cases, and its triggering conditions and reproducibility remain open questions. In addition, this reading is at summary scope and does not include figures or full experimental detail, so the concrete implementation behind the reported numbers and settings should be confirmed in the original text and code.
