Activation-Spectrum Regularization for Grokking-Style Generalization
Abstract
Overparameterized networks can memorize a small algorithmic training set long before they generalize. We test whether compressing the spectrum of hidden-state covariance favors representations that reuse directions across examples. Likelihood-gated activation-spectrum regularization (ASR) applies a trace-based Rényi-2 proxy, spectral entropy, or ridge Log-Det only as predictive confidence grows. The Log-Det gradient weights an eigenvalue by the inverse of its ridge-shifted magnitude, and Sylvester's identity permits an exact computation through the smaller of the batch Gram and feature covariance matrices. ASR lowered 15 of 18 task–objective means and won 66 of 87 matched comparisons across the fixed six-task suite. It reduced the required data fraction by 7.37 percentage points on average (95% stratified bootstrap interval [5.71, 9.27]), or 11.40% relative to the matched baseline ([5.87, 17.03]%). Rényi-2 lowered the mean threshold on all six tasks, while spectral entropy gave the lowest thresholds on two heterogeneous arithmetic mixtures. On UCI Superconductivity, every ASR objective improved all six mean test metrics under validation-selected and fixed-budget evaluation, although no objective won every metric. Full-run ModAdd trajectories show recurrent spectral-tail contractions aligned with transient accuracy drops. ASR is therefore a measurable representation-level bias whose benefit depends on architecture, task, and spectral objective.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.