acceptodds
Under review as a conference paper at ICLR 2027

Correction of Decoupled Weight Decay: Disentangle effective learning rate and weight norm control

Abstract

Since AdamW, decoupled weight decay has replaced regularization as the standard weight control method. More recently, steady-state analysis has suggested that decoupled weight decay should be proportional to learning rate squared to maintain stable weight norm, but such analysis can be derived from either orthogonality or independence of the weight and the gradient. We show that when applied to Muon / Scion, weight-gradient independence implies that the effective learning rate is momentum-dependent, a result that would be hard to explain with weight-gradient orthogonality alone. We show that optimal effective learning rate transfers across momentum values and enables log-time momentum scheduling for improved LLM training. Furthermore, models exhibit aggregated near-scale-invariance which explains the performance difference and altered model training dynamics due to weight decay correction. We conclude with both practical recommendations and implications of our results.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.