Demystifying Maximum Attention-Logit Dynamics in Muon Training
Abstract
Muon has been widely adopted in language-model training due to its superior efficiency, yet the origin of its excess late-stage attention-logit growth remains unclear. Our work systematically investigates the maximum attention-logit dynamics in Muon training. First, by comparing different attention implementations and numerical precisions, we show that excess late-stage growth appears only with FlashAttention (FA) in Brain Floating Point 16 (BF16), indicating that it is not intrinsic to the optimizer design but originates from numerical errors. We further find that, after an early transient, the resulting gradient errors become persistently rank-one, with stable directions and consistent signs, causing them to accumulate coherently rather than cancel and producing structured high-logit heads. Second, motivated by this structure, we analyze a representative retrieval task with additive rank-one gradient errors and prove that the precision-induced logit gap is negligible in the early stage but becomes dominant later, directly linking the observed errors to the two-stage logit growth. Finally, we identify the FA intermediate quantities most responsible for this growth and mitigate it by selectively increasing their precision. Experiments show that these interventions provide a new trade-off between training performance and logit magnitude relative to existing methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.