acceptodds
Under review as a conference paper at ICLR 2027

MALT: CURVATURE-AIDED MUON VIA LIGHTWEIGHT DIAGONAL PRECONDITIONING

Abstract

Muon has recently emerged as a promising alternative to AdamW for pretraining language models. By orthogonalizing momentum matrices using Newton–Schulz iterations, Muon mitigates the influence of gradient anisotropy by distributing the update more evenly across all singular directions, and enjoys an improved learning efficiency empirically. However, this operation does not explicitly account for the curvature of the loss landscape, and Muon may therefore still sensitive to curvature anisotropy. We bridge this gap by proposing MALT(Muon Augmented by Lightweight Two-sided preconditioning), which uses lightweight diagonal preconditioners to reduce the influence of curvature anisotropy to Muon. Specifically, MALT constructs two-sided diagonal preconditioners with low memory and computational overhead to capture a diagonal approximation of the curvature of the loss landscape. It then orthogonalizes the diagonally preconditioned momentum using Newton–Schulz iterations and maps the result back to define the update direction, while norm grafting controls the update magnitude. Convergence guarantees are provided for MALT in the stochastic non-convex setting. To improve its robustness to stochastic gradient noise, we further proposed MALTER (MALT with noise Adaptive stEpsize Rescaling). Experiments on GPT-2 Small, Medium, and Large pretraining show that the proposed methods outperform Muon while maintaining nearly the same memory footprint and wall-clock time.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.