Can Scaling Undo AI-Induced Data Pollution?
Abstract
AI models are increasingly deployed while their scale increases over time, and deployment can change the evidence available to train future models. We model this interaction through coupled dynamics of AI adoption and evidence dilution, while learner scale follows an exogenous path. We study whether increasing model scale can undo degradation created by earlier deployment, and how the answer depends on when scaling occurs and on curation. At a fixed learner scale, we show that sufficiently high curation leaves only the clean equilibrium, while sufficiently low curation can also sustain a stable equilibrium with persistent degradation. We then prove that there are settings in which, from the same initial state and at the same curation rate, starting at a higher scale can lead to the clean equilibrium, while starting at a lower scale and switching to the same higher scale after sufficient time can instead lead to persistent degradation. We show numerically that the qualitative behavior persists under alternative specifications of how evidence degradation affects model quality and future production. Separately, controlled fine-tuning experiments with Qwen2.5-1.5B show that degradation propagates through the evolving evidence even when each generation starts from the same pretrained checkpoint, while curation substantially reduces the effect. Thus, the effect of scaling depends not only on the eventual model scale reached, but also on the evidence produced along the way and on the strength of curation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.