acceptodds
Under review as a conference paper at ICLR 2027

Internal Data Repetition Destroys Language Models

Abstract

Language models are running out of high-quality training data, and even aggressively deduplicated corpora retain some amount of repetition. Earlier controlled studies predated Chinchilla-style scaling laws and could only measure the cost of repetition indirectly. We revisit repetition in the Chinchilla era, using a fitted no-repetition scaling law to report Compute-Equivalent Gain and Compute-Equivalent Loss. We show that under this modernized paradigm, repetition damage is systematic in three ways. First, when holding the amount of compute allocated to repeated data constant, eval loss peaks at an intermediate repeat count R; repeating a moderately sized subset a moderate number of times damages performance more than repeating a large subset a few times or a small subset many times. Second, the peak location is well described by a power law in model size; within the range studied, this trend shows that the most damaging repeated-data pool grows faster than the total training-token budget at a fixed overtraining multiplier. Finally, when repeated documents consume 10% of the FLOPs budget in a controlled exact-document repetition setting, the compute-equivalent loss can be large: on FineWeb-Edu-Dedup, the fitted peak for a Qwen3-style 344M-parameter model at OT=1 matches the loss predicted by a no-repetition calibration using about 67% of the FLOPs. We demonstrate that these phenomena are not language model specific, and can be analytically understood in a simple statistical model: a misspecified linear regression with verbatim duplicates reproduces the same qualitative loss peak, quantifying how such peaks can arise from a statistical tradeoff between memorization and generalization. Our findings add precision to the study of duplication in language models, allowing practitioners to quantify the wasted compute incurred by the presence and repeat structure of duplicates in pretraining corpora.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.