When Does Training on Synthetic Data Provably Work?
Abstract
High-quality training data are a major bottleneck in machine learning, motivating increasing use of generated data. We study when a binary classifier trained on synthetic data transfers successfully to a real distribution, under a shared conditional label distribution. First, we derive finite-sample excess-error bounds when the real covariate distribution is absolutely continuous with respect to the synthetic distribution under a controlled density ratio. We also obtain a bound for Lipschitz hypothesis classes in terms of the Wasserstein distance between the two distributions. Conversely, when the real distribution contains mass outside the synthetic distribution and the hypothesis class is sufficiently expressive, synthetic risk alone cannot uniformly control real-world error: a classifier can be arbitrarily close to optimal on synthetic data while approaching worst-case error on the uncovered real region. We confirm our results experimentally in controlled low-dimensional settings as well as on an application-scale dataset: Using 32,200 images generated as counterparts to real ImageNet images from four state-of-the-art image generators, we show the existence of an adversarial classifier which exhibits up to 0.5% error rate on the synthetic while degenerating to 99.25% on the real domain. At the same time, across a diverse collection of ordinarily trained architectures, synthetic and real test errors remain positively correlated. These results distinguish the worst possible case from typical empirical behavior and suggest that models trained using synthetic data should be rigorously validated on real data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.