Grokking Can Be Steered: Selecting Desirable Solutions By Choice Of Learning Rule
Abstract
During grokking, a network transitions from memorization to generalization, yet what shapes the neural implementation (circuit) of the generalizing solution and specifically the number of participating neurons (its size) remains unclear. We find that the learning rule, the combination of credit assignment algorithm and optimizer, controls this: On modular addition, every learning rule we study that groks recovers the same Fourier-based algorithm, yet implemented by circuits that differ in size. Why? Experimentally, we find that equalizing the magnitude of credit assigned to each neuron is responsible, e.g., expanding the effective circuit size, measured from the outgoing weights, from 60% to 99% of the layer's neurons at similar training and test accuracy. We vary which parameters are normalized together, from the whole layer to individual neurons, and find that only normalization per neuron fully distributes the circuit. Credit magnitudes are usually set by the built-in normalization of standard optimizers: Nero, which normalizes per neuron, distributes the computation, whereas Adam and seven others, including RMSProp, Muon, and Shampoo, use other normalization schemes and retain small, concentrated circuits. Distributed circuits tolerate neuron deletion better, while concentrated circuits transfer about twice as fast to a task that reuses the same features. Whether a learning rule groks at all is visible after 2% of training: The rules that memorize retain larger weight components in directions that leave training outputs locally unchanged. Learning rule choice thus steers circuit size, trading robustness against transfer speed without changing the learned algorithm.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.