Beyond Input Interpolation: Isolating Mixup's First-Order Regularization Kernel
Abstract
Mixup couples input interpolation with stochastic-gradient regularization. We isolate a first-order regularization kernel through deliberate purification: retaining a specified gradient structure while removing the input-dependent interpolation residual. For target-affine prediction gradients, target interpolation is exact at source inputs; the removed response captures input nonlinearity at the common mixed target. Sombrero Loss realizes this kernel exactly on original inputs with paired targets, then extends to independent targets. For targets of uniform mean, its expected gradient matches label smoothing at an analytically determined strength, while intrinsic noise has an explicit, data-dependent covariance. Under a prediction model, the average conditional per-example covariance and model-averaged generalized Gauss–Newton matrix share two components, specifying curvature-matching conditions and a local mechanism that can favor flatter regions. Mean-matched ablations and spectral measurements support regularization beyond average smoothing. CIFAR and ImageNet experiments show competitive generalization without input interpolation and improved stability over tested mixing strengths. Detection, translation, and class-imbalanced classification demonstrate the construction's applicability and target-prior flexibility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.