acceptodds
Under review as a conference paper at ICLR 2027

Data Coverage Gates the Grokking Transition in Small Transformers

Abstract

Grokking (sudden generalization long after memorization) is well-studied in small transformers on modular arithmetic, but the role of training-data diversity in the transition is unclear. In this controlled setting we isolate unique-pair coverage (the fraction of the p² ordered pairs seen) from total sample count. Our central finding is that unique-pair coverage is a causal, hysteretic boundary for the formation, but not the maintenance, of the Fourier circuit that enables grokking. Raising coverage past the critical value c* mid-training triggers grokking (9/10 seeds vs. 0/10 control); cutting it below c* after the circuit forms leaves it intact (10/10); and far below c* the failed state is operatively absorbing over 500,000 steps. Where the boundary is: within a budget of 100,000 steps, c* ≈ 24% of unique pairs at the default hyperparameters (15.6%–34.9% across weight-decay and learning-rate settings), with c*(p) ∝ p^(−0.36) under the held-out criterion, an intermediate scaling of the required pair budget between the O(p log p) and O(p²) predictions. What crossing it changes: the representation stays diffuse below c* and collapses onto a sparse Fourier circuit only above it, a circuit-efficiency count whose predicted exponent falls inside the measured interval. What it is not: the threshold reproduces under the without-replacement epoch protocol of prior work, is not an artifact of repetition or frequency skew, and carries over within the family to a second task and a two-layer transformer of the same width. Because data-scaling experiments vary coverage and sample count together, dataset-size effects should be read with coverage held fixed.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.