Controlled Demolition Is All You Need for Grokking
Abstract
There are a great many conjectures about how grokking is reached and what its properties are; many of them are supported by experiments and yet sit poorly with one another. I introduce three new methods for producing grokking — artificial ricochets (controlled demolition: the learning rate is escalated until the network breaks, and the network is then left to recover), an attention learning-rate over-drive, and an overdrive of the FFN input matrix together with row-wise weight normalization. All three run on the same training loop, the same architecture and the same initialization, and are compared there with the classical method, weight decay; this makes the paths to grokking commensurable and separates the intrinsic properties of the phenomenon from artifacts of the method that produces it. I show that the network’s work of generalizing can remain invisible to the usual metrics and is nonetheless brought out by applying another of these methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.