Inert on Algebra, Disruptive In-Context: A Causal Ledger for Rank Collapse in Grokking
Abstract
Effective-rank collapse is one of the most reproducible structural correlates of grokking. We ask whether preventing it changes generalisation, using a rank pin on the token embedding that is magnitude-matched online to the force driving the natural collapse, and we run the same pin on both sides of the comparison. On modular arithmetic the collapse is dispensable: pins at 0.9 times initialisation leave grokking on time under two pinned rank estimators and under a weight-decay-free optimiser, and still generalise on the attention-only architecture our in-context tasks use. Capping the embedding's Fourier concentration only paces grokking, with a pacing threshold of 0.118 at a 60k-step horizon (profile interval 0.110 to 0.127). On two sibling in-context circuits the same pin has the opposite effect. On bursty in-context labelling (BICL) none of 5 pinned seeds generalises by 1.3M steps, 45 to 97 times each seed's own unpinned grokking step, with floors held; on 1-hop induction the circuit does not form in any of 8 pinned seeds, against 8 of 8 paired baselines (exact paired p = 0.004). Each circuit has its own stable-rank threshold, 0.600 at 200k steps (bootstrap 0.550 to 0.666) and 0.845 at 75k (profile 0.804 to 0.895), and neither pre-registered normalisation carries one to the other, so the threshold is task-local. The block is not carried by the value of the global rank statistic, since matched random pushes keep rank near initialisation and generalise; it tracks how firmly the constraint binds. Four pre-registered mechanism attempts returned negative, so the mechanism remains open. Under standard softmax cross-entropy, weight decay drives the post-memorisation trajectory, and more than a quarter of the collapse happens after generalisation. Finally, the same statistic is timing-blind when read from the input embedding and tracks the grokking step when read from the unembedding, which explains how the correlate was misread.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.