acceptodds
Under review as a conference paper at ICLR 2027

Model Collapse Needs Only One Generation of Self-Training to Become Measurable

Abstract

Recursive self-training on synthetic data generated by antecedent models introduces severe distribution shifts and cumulative degradation, commonly referred to as model collapse. While prior literature debates the asymptotic limits of multi-generation feedback loops, the precise timing and measurable emergence of collapse remain underspecified. We present a controlled empirical and theoretical analysis demonstrating that model collapse is not an asymptotic artifact; it becomes strictly measurable after exactly one generation of self-training. Using discrete, countable state spaces and controlled MNIST image distributions, we track the total-variation distance between original and self-trained data generators. Our findings reveal that the initial collapse transition is highly pronounced yet sublinear across subsequent iterations, and downstream classification accuracy directly tracks this distributional degradation rather than masking it. Furthermore, we evaluate mitigation strategies, showing that data mixing ratios must scale logarithmically with generation depth to preserve support coverage. We conclude that model collapse constitutes an immediate operational bottleneck rather than a distant theoretical failure mode, necessitating strict provenance tracking in synthetic data pipelines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.