Group-Resolved Weight Decay Controls Grokking Arrival Across Checkpoint States
Abstract
Does grokking depend only on how much weight decay is applied, or also on where it is applied? We study two-layer Transformers using paired continuations from complete checkpoint states. We allocate decay between internal computation parameters (attention, MLP, LayerNorm; GC) and input/output parameters (embedding, head; EH), matching initial decay-only squared-norm contraction. On modular addition, shifting this matched total from EH toward GC advances sustained high test accuracy by steps on average, with the same direction in all 12 parents. Both groups accelerate arrival, but GC has the larger response in the tested dose units. At early checkpoints, high-GC continuations arrive in cases within steps, versus under low GC. Extending all 128 low- and intermediate-GC arms to steps yields arrival in every arm: the short-budget gap reflects delay for these runs, not permanent failure. The fixed allocation swap also advances arrival on permutation composition. Separate parameter-source interventions and a one-step response assay characterize complementary effects of inherited state, without establishing a mechanism for arrival. These results identify decay allocation as a control of grokking timing and show why finite-budget completion must be distinguished from eventual generalization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.