acceptodds
Under review as a conference paper at ICLR 2027

Grokking through the Lens of Minimum-Norm Interpolation

Abstract

Grokking reveals that fitting the training data and learning the underlying signal can occur at very different stages. However, existing theories offer limited quantitative insight into how this separation depends on inductive bias and signal structure. Our work addresses this gap by developing a statistical theory that quantifies how regularization geometry and signal sparsity govern generalization near interpolation. In particular, we focus on the prototypical setting of high-dimensional regression and identify regimes in which sparsity-promoting regularization makes exact interpolation much more accurate than approximate fitting. In strongly overparameterized noiseless problems, we prove a zero-one generalization law and construct a convex family of norms whose interpolators transition from the trivial risk of the all-zero predictor to exact recovery, while keeping the training error always equal to . In the regime where feature dimension and sample size scale proportionally, we provide a precise characterization of the training and generalization errors along -regularization paths. This in turn allows us to quantify the generalization gain that remains near interpolation: we show that this gain increases as the norm becomes more sparsity-promoting and as the target becomes sparser, with a sharp drop in generalization reached for noiseless data and regularization. Experiments on diagonal linear networks and transformers trained on modular arithmetic demonstrate that the prediction of our theory extend well beyond the linear setting of the analysis.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.