Mixup Mitigates Instance-Level Overfitting under Structural Overparameterization
Abstract
Overparameterization is commonly measured by the number of fitted model parameters relative to the training sample size. In this paper, we examine mixup augmentation through the complementary notion of *structural overparameterization*: a nonlinear student contains directions that are unnecessary for a linear teacher but can fit observation noise. For Gaussian covariates, we derive precise high-dimensional characterizations of the prediction risk under full-pair mixup. Our analysis combines a decomposition of pairwise interactions with leave-many-out arguments to handle the complex dependence among mixup samples generated from training data. The obtained risk formulas separate background prediction error from label-noise sensitivity, revealing how mixup changes their tradeoff under *homogeneous* and *heterogeneous* noise. In particular, under explicit conditions, we show that mixup achieves *strictly lower risk* than ordinary nonlinear regressor, with the ridge penalty optimized separately for each method. This advantage consequently extends to mixup jointly tuned over its mixing weight and ridge strength. Numerical experiments support the predicted risk comparisons. Motivated by this analysis, we propose a selective mixup rule based on training-time fitting progress in neural networks with learned features on three real-world regression datasets. Under training-label perturbations, selected mixup reduces test error relative to empirical risk minimization and separately tuned weight decay.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.