acceptodds
Under review as a conference paper at ICLR 2027

Future Data Allocation Shapes Repetition-History Persistence in Language Model Pretraining

Abstract

High-quality pretraining data are finite, so repeated exposure is increasingly common. Prior work has mainly examined how repetition affects the final model. We instead ask whether predictive differences caused only by when repeated data appear persist under later training. We train matched model pairs on the same document-occurrence multiset using different repetition schedules, then continue both models on identical ordered batches. Under a fixed continuation budget, the remaining difference changes systematically with support–reuse allocation: concentrating tokens on repeated visits to narrow support can increase disagreement, whereas distributing them over broader support produces greater contraction. Token count alone does not explain this response. At 124M parameters, 15M tokens over broader support produce more contraction than 30M tokens of restricted-support reuse. A shared-first-update intervention localizes the same behavior: after an identical first update, every tested revisit increases disagreement, while matched distinct batches yield a larger one-step decrease. The allocation contrast replicates in a document-disjoint FineWeb-Edu construction and on C4, and remains positive across model scales and optimizer controls. These results show that the fate of a repetition-induced predictive difference depends on both the amount of continued pretraining and the allocation of its future data.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.