acceptodds
Under review as a conference paper at ICLR 2027

ADAMW MAINTAINS WHAT DELAYS GROKKING

Abstract

A network that groks fits its training data thousands of updates before it generalizes. What keeps it waiting? We find that the wait depends on which content the network's embedding holds once the training data are fit, and that the AdamW updates that keep rebuilding this content against weight decay lengthen it. In one- and two-layer Transformers learning modular arithmetic, the operand embedding decomposes into Fourier frequencies, so we can fork a run and change one thing: delete what a set of frequencies holds, put it back, or stop the optimizer from rebuilding it. Masking a random half of the frequencies during training halves the time to generalize, although the normally trained model relies on that half, so removing what a trained model relies on need not slow its learning. In one-layer models, putting back what those frequencies held before generalization, the network's own early content, into the same frequencies of the same model state delays generalization by over two thousand updates; a scrambled copy of equal energy, the network's later content, or another network's early content barely delays it. Weight decay alone would erase most of this early content within 2,000 updates, yet in ordinary training it keeps its size: AdamW's updates put back what decay removes. After the same deletion, blocking the rebuild brings generalization forward beyond deletion alone and beyond a control that shrinks each update to the norm blocking would leave. Over that control, the paired median gain is 850 updates at one layer and 800 at two, both on fresh seeds. Deleting the content and blocking its rebuild, which acts on a single direction of the embedding, recovers about 0.9 of the speed-up from masking. Part of the wait thus lies in early embedding content that AdamW keeps in place.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.