Zero-Data Self-Play Pretraining: a generator and a learner trained in tandem from random initialization show zero-shot loss scaling predictably with self-play compute across several natural datasets
Synopsis
The work introduces Self-Play Pretraining with Zero Data as an initial proof-of-concept: starting from random initialization, a generator proposes programs interpreted by a universal Turing machine to produce byte sequences while a learner autoregressively predicts those byte sequences with standard cross-entropy, and the generator is trained with reinforcement learning to produce sequences at the frontier of the learner's capabilities, yielding an adaptive curriculum; because neither generator nor learner is trained on natural data, zero-shot performance on natural data serves as a clean test of transfer, and across several natural datasets zero-shot loss exhibits predictable scaling in compute, with the models also exhibiting in-context learning and discovering recognizable mathematical
Figure 1: Self-Play Pretraining with Zero Data . Starting from randomly initialized models and using only synthetically generated data, self-play produces predictable scaling on held-out natural data. (Top) Our procedure casts synthetic data generation as search over the space of computable structure. A generator proposes programs, which are executed to produce byte sequences used to train a learner by next-token prediction. The generator is trained via reinforcement learning to propose programs near the frontier of the learner’s capabilities, measured by how strongly the learner’s gradients align with the learner’s recent learning trajectory with respect to an AdamW-preconditioned inner product. (Lower left) Self-play produces predictable improvements in validation loss with compute across natural datasets. We show the compute-optimal frontier over model sizes, number of self-play rounds, and ensemble sizes. (Lower right) The resulting learner exhibits in-context learning on held-out tasks without any gradient updates. We report empirical success under greedy decoding as a function of the number of in-context examples m m , averaged over independently sampled task instances.
arXivInterpretation
It proposes and implements a pretraining procedure that casts synthetic data generation as a search over the space of all computable structure, taking inspiration from Solomonoff induction. Whereas pretraining data has largely been curated on the model's behalf, this work lets the model generate the data most useful for its own improvement, in principle shifting the source of training data from human knowledge to compute. The text defines the search space via a universal Turing machine interpreting the generator's proposed programs and states that this imposes little domain-specific structure; this is a method-level design statement within a proof-of-concept.
It builds a two-model tandem training mechanism from random initialization: the learner predicts byte sequences autoregressively with standard cross-entropy, while the generator is trained with reinforcement learning to produce sequences at the frontier of the learner's capabilities. The generator is not a fixed data distribution but is rewarded for approaching the learner's current capability boundary, forming an adaptive curriculum that changes as learning progresses. The text explicitly gives both training objectives (cross-entropy and reinforcement learning) and the generator's frontier-seeking target, but the summary provides no algorithmic details, hyperparameters, or scale.
Across several natural datasets, zero-shot loss exhibits predictable scaling in compute. Because neither generator nor learner is trained on natural data, this zero-shot performance is a clean test of transfer, extending scaling behavior from natural-data pretraining to self-generated data. The text reports that 'Across several natural datasets, zero-shot loss exhibits predictable scaling in compute', but the summary gives no dataset names, loss values, or fitted form.
During training the models exhibit in-context learning and discover recognizable mathematical sequences. These behaviors emerge without any natural-data supervision, purely from self-play over generated byte sequences, suggesting the procedure yields behavior beyond fitting the generated distribution. The text states that the models 'exhibit in-context learning' and 'discover recognizable mathematical sequences'; these are qualitative observations, and the summary offers no quantitative metrics or examples.
Perspective
The result speaks to readers interested in the feasibility of pretraining on self-generated data, especially those asking whether the data source can shift from human knowledge to compute. It applies under the setting in which both generator and learner start from random initialization and neither is trained on natural data, with zero-shot loss on natural data as a function of self-play compute as the evaluation. Within that setting, it shows that zero-shot loss can exhibit predictable scaling, accompanied by in-context learning and the discovery of mathematical sequences. Moving forward, the same framework could be examined at larger compute and across more natural datasets, and the interaction between the generator and the learner's capability frontier could be traced as training evolves.
The summary does not give the names or number of natural datasets, the zero-shot loss values, or the form of the scaling fit, nor does it state model scale, compute range, or training duration, so the concrete shape of 'predictable scaling' still requires the body. In-context learning and 'recognizable mathematical sequences' are qualitative statements in the summary, without quantitative metrics or examples. In addition, the work describes itself as an 'initial proof-of-concept', so how far its conclusions hold at larger compute and larger models is an open question. This summary is based on the abstract only, without figures or experimental details, and those gaps may affect judgment of result strength.
