acceptodds
Under review as a conference paper at ICLR 2027

Model Collapse Is Avoidable Even as Real Data Vanishes: A Sharp Criterion for Iterative Gaussian Estimation

Abstract

Large generative models are commonly trained on data collected from the open web, which may contain data produced by previous generations of the same model. Previous work has shown that recursively training on model-generated data can cause a progressive degradation in model performance, a phenomenon known as *model collapse*. Yet, currently, generative models continue to improve dramatically over time. Existing theoretical research suggests that this is enabled by the availability of new real data at each iteration, kept separate from model-generated data. However, exact separation is not realistic, and uncurated data is still commonly used. We bridge this gap by considering a less restrictive and more realistic setup in which real and generated data are not handled differently during training. Specifically, we present a comprehensive theoretical study of iterative estimation of the mean and covariance of a multivariate Gaussian data distribution. At each iteration, samples are drawn from a mixture of the real data distribution and the estimated distribution from the previous iteration, but are used indistinguishably. We provide a single equivalent condition for consistency of the mean and covariance estimators in Euclidean norms, and convergence to zero of the expected squared -Wasserstein distance between the estimated and true distributions. Using this result, we then show that consistency holds, i.e., model collapse is avoided, for any fixed real data fraction above zero, and even for a fraction that vanishes across iterations, provided it does not vanish too quickly. However, more real data (or data curation) remains important at practical finite times: we show that the convergence rate of the mean estimator deteriorates as the real data fraction approaches zero. We support our results with an empirical investigation of the theoretical setup, and complement our theoretical work with experiments on iterative training of a diffusion model on CIFAR-10 that correlate with our theory.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.