Scale Dependent Data Duplication
Abstract
Data duplication during pretraining can degrade generalization and lead to memorization, motivating aggressive deduplication pipelines. However, at web scale, semantically equivalent documents (e.g. translations) may induce redundant training signals once models become sufficiently capable. We present evidence that duplication is scale-dependent in two ways. First, as model capability increases, cross-entropy loss gradients for semantically equivalent documents become more aligned. Smaller models produce gradients that reflect surface similarity rather than semantic similarity. Second, we embedded all 192 million FineWeb-Edu-Dedup documents using EmbeddingGemma-300m. For moderate corpus sizes, nearest-neighbor cosine gap follows an isotropic power law baseline, but deviates sharply as corpus size grows to hundreds of billions of tokens, indicating accelerated semantic collisions. Finally, controlled pretraining on data sampled with replacement from finite pools of unique documents shows mild degradation for small models but rapidly increasing loss penalties for larger models, breaking naive scaling extrapolation. We derive scaling laws that account for limited semantic uniqueness, allowing for more accurate prediction at scale.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.