Benign Misfitting Beyond One Spike
Abstract
Benign misfitting occurs when accurate prediction requires larger empirical training loss than the zero predictor's. We show this phenomenon occurs for noiseless overparameterized Gaussian regression with signal spikes of unequal strengths and independent training and test pairs from one distribution. Specifically, we show that there exists a range of sample sizes in which every sufficiently accurate linear predictor in the span of the training data must misfit. So good generalization here cannot be explained by good training fit. A reflection argument extends this limitation to all learned linear predictors when no signal spike orientation is preferred in advance. We then characterize when a single epoch of zero-initialized stochastic gradient descent can reach predictors that generalize well, despite this requiring actually ascending in empirical training loss. An explicit pair of examples is given to separate the geometry forcing misfitting from whether such misfitting predictors are learnable by a single epoch of constant-rate SGD.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.