acceptodds
Under review as a conference paper at ICLR 2027

Rewriting Scaling Laws for Synthetic Data with Compute and Data Constraints

Abstract

Synthetic rewrites, where a language model rewrites a natural document into a target style, are now a standard ingredient of pretraining corpora. Yet critical aspects of their use remain poorly understood, such as how much of the training data should be rewritten, given that generating rewrites costs compute that could instead train the model, and whether a model trained on rewrites can outperform the rewriter. We study these two questions by proposing a rewriting scaling law that predicts loss from model size, training tokens, the fraction of training tokens that are rewrites, and the number of epochs over the natural data, charging the compute spent generating rewrites. We fit it in two settings. In the compute-limited setting, natural data is plentiful and rewrites come from a third-party rewriter trained with more compute than our largest run. We find that the estimated optimal rewrite fraction increases with compute, from under at FLOPs to at FLOPs. Our scaling forecast extrapolates in FLOPs to train a 22B-parameter model, whose final loss is within of our prediction. This model achieves a better BPB than its rewriter on several web domains at one-tenth the rewriter's training compute. In the data-limited setting, natural data is a fixed seed set and the rewriter is a model of the same size as the target, pretrained on that seed set. Again the optimal rewrite fraction increases with the compute budget. We also demonstrate that a model trained on rewrites of its own pretraining data, generated by a rewriter of the same size, beats the best run on natural data alone at every model size up to 3B. Our results show that synthetic rewriting can be studied through the scaling law lens: that optimal methodological choices are predictable and scaling does not necessarily break down beyond the rewriter's capabilities.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.