acceptodds
Under review as a conference paper at ICLR 2027

The Statistical Origins of Model Collapse: Resampling, Retention, and Exact Variance Recursions

Abstract

Model collapse — the degradation of generative models trained recursively on their own output shumailov2024ai — is usually explained by properties of the neural networks involved. We show the dominant mechanism is instead statistical: plain kernel density estimation (KDE), repeatedly refit to its own samples, reproduces qualitative collapse with no neural network involved. Using this minimal setting, we adjudicate between two follow-up results that pull in opposite directions — gerstgrasser2024collapse's claim that accumulating synthetic data avoids collapse, and dohmatob2024strong's claim that no synthetic fraction is safe — and show that retention of past data, not the injection rate, is what separates convergence from divergence. This retention effect survives sweeps over dimensionality, tail weight, and seven estimator families — including a normalizing flow, a GRU, and a causal self-attention Transformer — and reproduces on real images and real text-like sequences. We then derive exact variance recursions for a tractable Gaussian-MLE idealization of both regimes, verified against Monte Carlo simulation, and test these predictions directly against a small VAE trained by gradient descent: the accumulation-regime prediction matches closely, while under replacement the neural estimator collapses faster than the idealized statistical mechanism's own worst case. Neural networks, in other words, do not create model collapse so much as amplify a mechanism that operates identically in their complete absence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.