acceptodds
Under review as a conference paper at ICLR 2027

Demystifying Maximum Attention-Logit Dynamics in Muon Training

Abstract

Muon has been widely adopted in language-model training due to its superior efficiency, yet the origin of its excess late-stage attention-logit growth remains unclear. Our work systematically investigates the maximum attention-logit dynamics in Muon training. First, by comparing different attention implementations and numerical precisions, we show that excess late-stage growth appears only with FlashAttention (FA) in Brain Floating Point 16 (BF16), indicating that it is not intrinsic to the optimizer design but originates from numerical errors. We further find that, after an early transient, the resulting gradient errors become persistently rank-one, with stable directions and consistent signs, causing them to accumulate coherently rather than cancel and producing structured high-logit heads. Second, motivated by this structure, we analyze a representative retrieval task with additive rank-one gradient errors and prove that the precision-induced logit gap is negligible in the early stage but becomes dominant later, directly linking the observed errors to the two-stage logit growth. Finally, we identify the FA intermediate quantities most responsible for this growth and mitigate it by selectively increasing their precision. Experiments show that these interventions provide a new trade-off between training performance and logit magnitude relative to existing methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.