Detecting overfitting in Neural Networks during long-horizon grokking using Random Matrix Theory
Abstract
Training neural networks without overfitting is difficult; detecting that overfitting is difficult as well. We present a data-free Random Matrix Theory method for detecting the onset of overfitting from the trained weights alone. For each layer, we shuffle the weight matrix element-wise, , fit the shuffled matrix's empirical spectral distribution with a Marchenko–Pastur law, and identify large outliers beyond the finite-size right edge. We call these outliers Correlation Traps. In long-horizon grokking, traps form and grow in number and scale during anti-grokking, a post-generalization phase in which training accuracy remains high while test accuracy declines. Across an MNIST MLP, a modular-addition transformer, and a GPT2-style composition model, Correlation Traps distinguish anti-grokking from pre-grokking despite similar train and test behavior. Following Li and Sonthalia's spiked-regression account, a separated spike may be benign or harmful depending on its alignment with the learned function. We estimate this alignment without task data by passing random inputs through the trained model and measuring the Jensen–Shannon divergence between output logit distributions before and after suppressing each detected mode. The resulting scores recover the behavioral ranking obtained with in-distribution inputs, while continued suppression of the detected modes improves generalization more than an equal-count, equal-update-energy largest-weight edit. We also find Correlation Traps in foundation-scale language models, where their effect on generalization remains to be tested. These results identify randomized weight spectra as a data-free signature of anti-grokking and a practical screen for potential harmful overfitting.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.