From Anti-Generalization to Grokking: Train–Test Representation Kernel Dynamics in Modular Addition
Abstract
Grokking is a phenomenon in which a model generalizes long after fitting the training data. It is commonly attributed to delayed feature learning, yet a quantitative explanation linking evolving hidden representations to generalization remains incomplete. We develop a quantitative framework based on evolving train–train and train–test representation kernels, which directly connect hidden representations to training and test logits. An approximately label-aligned residual persists after fitting and sustains slow feature learning under weight decay. Within this framework, we show that early random-feature structure actively drives anti-generalization, pushing test accuracy from chance to nearly zero through spurious train–test correlations, whereas subsequent feature learning reorganizes the cross-kernel toward task-aligned structure and drives generalization. The sharp transition is controlled by finite-width fluctuations associated with the effective number of feature-aligned neurons. At late-time equilibrium, the learned feature-kernel geometry further predicts the cross-entropy loss through a mean-field scaling law in weight decay and problem size. Together, these results provide a quantitative representation-level account of grokking from early anti-generalization to late generalization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.