Persistent Overfitting: The Statistical Price of Interpolation
Abstract
What is the statistical price of fitting noiseless training examples exactly? In noiseless linear regression with sufficient dimension, requiring interpolation raises the minimax expected population risk from to , with the same data and linear predictor class. Gradient descent and fixed-dataset SGD incur this statistical price at convergence, matching the minimax rate among interpolating estimators. Yet, on a single instance, their -point prefix averages achieve expected population risk at , followed by a persistent deterioration to , even though the iterates never move farther from the target in Euclidean distance. We explain this behavior through shared constraints and missing coordinate penalties. For general realizable quadratics, we sharply characterize how the interpolation price depends on dimension, population trace, and sample rank. Finally, we study how accurately the risk of the converged predictor can be certified from observed data. With fresh validation examples, the optimal certification scale is , even when the true risk is zero.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.