Upcycling the Low-Quality Corners of the Web for Pretraining
Abstract
Modern language model pretraining pipelines collect far more web text than they ultimately use, discarding large fractions through quality filtering. We call this **token waste**: potentially useful tokens are acquired during data collection but removed before training. Existing synthetic data methods mostly *recycle* data by rewriting high-quality documents that already survive filtering. We instead ask whether data deemed too low-quality for training must actually be discarded; can they instead be **upcycled** by transforming low-quality documents into useful pretraining data. Concretely, we train a small language model with preference optimization over learned quality and influence-based utility signals, then use it to upcycle the lowest-quality 14B tokens from C4 (i.e., wasted tokens). Across language models from 19M to 1.4B parameters, upcycled tokens consistently outperform wasted tokens on in-domain and out-of-distribution language modeling, as well as downstream performance. When combined with high-quality data, upcycled tokens recover much of the benefit of adding more high-quality data. Additional experiments show that upcycling preserves factual knowledge lost through filtering, improves existing synthetic data recipes, remains effective with substantially smaller synthesizers, and extends over-training gains beyond what is achieved by repeating high-quality data. Together these results show that **token waste can be prevented,** with benefits extending well beyond recovering pretraining utility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.