acceptodds
Under review as a conference paper at ICLR 2027

Representability Decides Whether Grokking Can Occur

Abstract

A flat test curve has several possible causes. The target may lie outside the architecture's function class, or the training sample may permit memorisation without constraining the population solution, or the optimiser may need a long time to find a generalising representation. Learning curves do not tell these apart, and existing grokking analyses begin only once memorisation is feasible. We separate them before training. Our instrument is an exactly solvable two-layer network with holomorphic activation on roots-of-unity-encoded modular inputs. Its expressible class is an algebraically fixed -dimensional subspace of characters of , independent of width, and a task is representable exactly when its discrete Fourier support lies in . For linear-phase targets this becomes the arithmetic criterion , an identity in on the canonical residues rather than a congruence modulo . From the class we extract two quantities computable in advance, the spectral representability defect and the conditioning of the class restricted to the training sample, and prove two consequences that do not follow from the definitions. The global obstruction reaches the finite sample as a width-independent training-loss floor with an explicit rate, so a non-representable target cannot be memorised even though a flexible model would interpolate it. For representable targets under a well-conditioned restriction, training and population losses are sandwiched together at every iterate, so the prolonged separation that defines grokking is structurally excluded. Across 585 runs the pre-training prediction matches the observed regime with accuracy. Two experiments carry the diagnostic beyond the solvable model. A bottlenecked one-hot ReLU network, given no Fourier basis, passes through four regimes as the bottleneck widens. At test accuracy stops at a ceiling that a fourfold step budget does not lift, and at the outcome stops being a function of width: six seeds of eight escape after delays spanning two orders of magnitude, and two never reach the chance rate at all. Swapping AdamW for vanilla gradient descent at fixed model and target reverses the degree-dependence of the acquisition delay in all sixteen cells of a four-target grid crossing vanilla gradient descent with three AdamW weight decays, so an observed scaling of grokking time can be optimiser-induced and not a property of the landscape. Which learning regime exists at all is set by the geometry of the effective function class, not by parameter count.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.