Why Does Rewriting Pretraining Data Improve Generalization?
Abstract
Language models are increasingly pretrained on synthetic rewrites of human text, which can improve downstream generalization. This gain is puzzling, as rewriting models can benefit substantially larger learners even on capabilities they do not themselves possess. We present an explanation through a rewriting-pretraining gap. We show that regularities in natural language allow rewriting models to learn simple transformations that generalize to unseen inputs. Applying these transformations changes the prediction problems encountered during pretraining, increasing the computationally bounded information accessible to the learner and inducing new representations useful for downstream tasks. We formalize this mechanism using algorithmic information theory and validate its predictions in controlled experiments. Our account further predicts that rewriting need not preserve factual content. Consistent with this prediction, we show that fictional rewrites improve compositional generalization on a challenging two-hop knowledge task.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.