Does U-W Ratio Alignment Benefit Muon in LLM Pretraining?
Abstract
Most top-performing entries in the `modded-nanogpt` optimization benchmark use a "u/w floor". The u/w floor lower-bounds the Frobenius-norm ratio of each block's pre-learning-rate update to its weight (*u-w ratio*). Like decoupled weight decay, the u/w floor prevents updates from becoming negligible relative to the current weights during training. Unlike decoupled weight decay, however, the u/w floor also aligns u-w ratios across blocks. Specifically, it aligns the u-w ratios of most weight blocks in a model architecture to a common value on most training steps. We therefore ask whether this u-w ratio alignment is an underlying mechanism of the u/w floor's empirical advantage. To study this question, we remove the u/w floor and directly enforce u-w ratio alignment on Muon updates. This gives *MuonL*, which assigns a common ratio to all Muon-updated blocks, and *typed MuonL (t-MuonL)*, which aligns only blocks with the same Transformer role. To assess these policies independently of particular learning-rate values, we derive a measurable local diagnostic from a theoretical model of one-step loss decrease. When all Muon-updated blocks share the same learning rate, the analysis yields a covariance-based sufficient condition for when MuonL improves upon Muon. When the learning rates can vary across matrix types, it yields a generalized-eigenvalue criterion for when t-MuonL outperforms Muon and vice versa. We then evaluate the diagnostic at checkpoints along Muon training trajectories, using each checkpoint as a starting point to compare Muon's one-step decrease with the counterfactual decreases under MuonL and t-MuonL. It favors t-MuonL over Muon in both cases and MuonL over Muon in the former. Controlled runs using the same training configurations corroborate the diagnostic, with lower final validation losses for MuonL or t-MuonL than for Muon. These results support u-w ratio alignment as a contributing factor to the u/w floor's empirical advantage and a useful intervention for Muon-based LLM pretraining.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.