Self-Play Pretraining with Zero-Data
Abstract
Advances in language modeling have been driven by scaling pretraining on ever more data. Yet, training data is still largely curated on the model's behalf. A more general approach to pretraining would let the model learn to generate the data that is most useful for its own improvement. This would provide an effectively unbounded source of training data, relying on compute rather than the total sum of human ingenuity. We introduce **Self-Play Pretraining with Zero Data**, an initial proof-of-concept towards this vision. Starting from random initialization, two autoregressive transformers learn in tandem: a generator proposes programs interpreted by a universal Turing machine, generating byte sequences, while a learner autoregressively predicts these byte sequences. The learner is trained with standard cross-entropy, while the generator is trained with reinforcement learning to produce sequences at the frontier of the learner's capabilities, yielding an adaptive curriculum. We use a universal Turing machine because any computable data-generating process can be represented as a program, providing a broad search space with minimal domain-specific assumptions. Self-play then learns which parts of this space produce useful training data, rather than requiring us to specify such structure in advance. We test whether zero-shot performance on natural data improves predictably with self-play compute; this is a clean test of transfer since neither generator nor learner is trained on natural data. Across several natural datasets, zero-shot loss exhibits predictable power-law scaling in compute, with scaling exponents comparable to those obtained by training directly on natural data. This suggests that learning universal predictive structure may play an unappreciated role in driving pretraining gains; we formalize this via a new scaling law ansatz. The resulting models also exhibit in-context learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.