What a Grokking Plateau Can Be Rescued By Is Not What It Is Spent Building: Component Surgery in a Glass-Box Network
Abstract
A grokking plateau poses two questions usually conflated: which component, supplied, ends it, and which one training spends it building. We separate them — exploitable from deferred — in a glass-box network whose hidden units all compute one fixed primitive, , and learn only routing and gains, so its component boundaries are architectural. It groks modular addition in 17/19 runs. Implanting a run's 100k-step embedding into its own 30k-step checkpoint reaches 0.9 test accuracy in 1.2k steps against 12.7k for a fresh-optimizer fine-tune-only baseline; a capacity-matched readout head is slower, a random direction at matched norm does not rescue on this task, and rescaled host structure ends below baseline at 20k. An untrained cosine bank rescues if its frequencies are position-shared. A registered control refuses the missing-component reading: the host already spans 87% of the donor embedding's norm, and that in-span part reads higher at 2k (0.823–0.834 against the orthogonal part's 0.289–0.527, three hosts, though it crosses 0.9 first on only two of them) — yet 72% of it is not the host's own structure rescaled. The orthogonal part rescues too, so the in-span part is better early, not necessary: amplification fits as well as supply. Either way the rescue works only in a host short of the solution. Exploitable and deferred come apart on both tasks — embedding exploitable, wiring deferred — but the wiring's exploitability inverts: harmful early on modular addition, immediately useful on parity. What parity's embedding defers is its allocation of power to the support bits, not its convergence. All results are small-scale, on two task families.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.