acceptodds
Under review as a conference paper at ICLR 2027

MultiNormStep: Leaving the Corner of Steepest Descent

Abstract

First-order optimizers are the workhorse of large language model training, and after a decade in which AdamW was the default, Muon has recently become the first optimizer to beat a well-tuned AdamW across model and data scales. Yet Muon acts only on the 2D matrix weights; the remaining parameters (the normalization scales, biases, and most consequentially, the token embedding and LM head) are left as a compromise to a default AdamW. This split looks like two unrelated optimizers bolted together, but through the lens of steepest descent the pieces are more related than they appear: signSGD and Muon are the -norm corner read in two gauges, the box on coordinates and the Schatten- ball on singular values, and Adam is in turn a signSGD that shrinks every coordinate strictly inside that box. SignSGD sits aggressively at the corner and Adam timidly within it, yet neither leaves it. Rather than commit to a single corner, we constrain a of norms through a Young function, obtaining (MN-step): a family with a closed-form coordinatewise update that both the single-norm methods (signSGD, PowerStep, normalized SGD) and new interior maps, a bounded and an unbounded . The identical construction lifts from coordinates to the spectrum, where it recovers Muon, so these maps and Muon are two faces of steepest-descent principle. Under full-parameter training our bounded sigmoid step stays on par with a carefully tuned AdamW across GPT-2 124M/350M and Qwen3 0.6B/4B; its decisive role is the group Muon leaves to AdamW, where, in a Muon-hybrid that fixes Newton-Schulz Muon on the 2D weights and swaps only that branch, the sigmoid step beats the AdamW branch on validation at Qwen3 0.6B and 4B and again outperforms AdamW at two scales of a proprietary mixture-of-experts model (Model-X). These consistent gains show that widening the geometry beyond a single norm genuinely helps, precisely on the parameters Muon leaves to AdamW.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.