acceptodds
Under review as a conference paper at ICLR 2027

Tuned Convex Augmentation Beats the Sample It Is Built From

Abstract

Data-augmentation methods such as SMOTE and mixup create synthetic points as random convex combinations of real examples. We ask when the resulting synthetic distribution is closer to the data distribution, in Wasserstein distance, than the real examples it is built from, whose empirical distribution converges at rate . Mixing nearby points smooths the sample. This reduces fluctuation at the cost of a bias of order , where is the neighbourhood size. For a smooth density bounded away from zero, we prove that -NN SMOTE with tuned beats for every in every dimension . Its rate reaches once , and for a non-flat density on the torus no local mixing scheme with a bias of order can do better. The gain carries over to a finite training set only if synthetic points make up most of it. In experiments, tuned mixing of pairs beats the real sample for as small as . In full dimension a tuned kernel does better, while on the two manifolds we test mixing beats an isotropic kernel at every sample size and ambient dimension.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.