acceptodds
Under review as a conference paper at ICLR 2027

PRECOCIAL: Pretraining with Code Curriculums via Iterative Search for Accelerated Learning

Abstract

Fully synthetic pretraining data generated by symbolic programs can bootstrap downstream training on naturalistic tasks, such as language modeling, but the gen- erating programs have so far been designed by hand. We introduce PRECOCIAL, a method for automatically optimizing programs that generate synthetic pretrain- ing data. Starting from an initial generator program, an LLM iteratively rewrites the program, and each rewrite is scored by pretraining a randomly-initialized transformer on its data and fine-tuning it on the downstream task. Because ev- ery score requires training, PRECOCIAL keeps a single network for the entire search, continuing its pretraining for each accepted program instead of training a new network per candidate, making it significantly more efficient. On four lan- guage and code datasets, at a fixed fine-tuning budget, searched programs reach lower perplexity than hand-designed curricula, published procedural generators, random program mutation and the LLM’s own first draft, and reach the same per- plexity as the hand-designed curricula with 9–33% fewer fine-tuning steps. We also observe that using a searched program generator from one domain transfers to datasets in other domains. Additionally, searched programs allow for better downstream performance on four different planning tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.