acceptodds
Under review as a conference paper at ICLR 2027

To Bias or Not to Bias: A Phase Transition in Learning from Synthetic Data

Abstract

Real-world dataset curation is expensive. Pre-trained machine learning models provide a powerful means of generating synthetic data to augment limited real data for training downstream models. However, while synthetic-data augmentation has demonstrated empirical success, there is little theoretical understanding of when synthetic data actually helps. Here, we tackle fundamental questions: when can a model trained on synthetic data outperform its teacher, how much synthetic data is required, and how should the student be trained? We study these questions through a student-teacher framework in which both the teacher learns the distribution from the real data, and the student learns from the synthetic data generated by the teacher. When the hypothesis class for both teacher and student is the family of Gaussian distributions, we show if both teacher and student are maximum likelihood estimators (MLE), then the student can never outperform the teacher in terms of Kullback-Leibler (KL) divergence to the true distribution. Crucially, allowing the student to use an appropriately biased estimator changes this picture. We characterize a phase transition in the number of real and synthetic samples: the boundary nearly scales *quadratically* with the number of real samples; below which no student that is scalar multiple of MLE can outperform the teacher, and above the boundary a suitably biased student can achieve strictly lower KL divergence to the true distribution and consequently improved generalization. This phase transition provides exact practical guidance on how to train student models on synthetic data. A natural question is whether we can repeatedly use the student model as a teacher to train a new student model, and so on, to further boost the performance of the student model. We show that achieving an improvement at every stage requires a sample complexity that grows exponentially with the length of the chain. Moreover, we show that a two-stage mechanism can achieve the same asymptotic performance as arbitrarily long chains, revealing a fundamental *No Free Bootstrapping* phenomenon. Finally, we demonstrate the applicability of our theoretical framework in both discriminative and generative models, including linear discriminant analysis and linear flow matching.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.