Generative Curation Yields Bigger and Better Data for Pretraining LLMs
Abstract
Modern large language models (LLMs) are pretrained largely on data curated from the internet. Because of the web's vast scale, pretraining relies on inexpensive curation: heuristic rules and classifiers to extract text from HTML and filter documents. We introduce **Generative Curation with LLMs** (GC-LLM), which replaces both stages with an LLM and a natural language specification of useful training data. The LLM follows the spec to transform raw HTML into training text or discard unsuitable documents. On a fixed pool of Common Crawl data, GC-LLM yields – as many training tokens as DCLM-baseline and Nemotron-CC. Across models with 157M, 1B, & 2.9B parameters and two LM evaluation suites, models pretrained on these tokens achieve –% lower bits per byte (bpb) than those trained on either baseline. To make this approach practical at web scale, we distill GC-LLM into a fast cascade comprising FastText, a 26M-param transformer, and a 68M-param BERT model, which we call **Generative Curation with Cascade** (GC-Cascade). GC-Cascade agrees with the LLM curator on % of filtering decisions and further reduces best observed bpb by –% on 1B and 2.9B models. We use this cascade to begin constructing Atlantic-CC, a projected 20T-token dataset, and release code and models to reproduce our data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.