acceptodds
Under review as a conference paper at ICLR 2027

MuonIO: Principled Norm-Aware Descent for Embedding Tables and Language Model Heads

Abstract

The Muon optimizer derives its update rule for hidden linear layers by solving a local linearization of the loss penalized by the spectral norm, motivated by an RMS-stability argument for dense linear layers. Standard Muon implementations, however, exclude the input (embedding table) and output (language model head) layers from this principled treatment, for which they use AdamW instead. We present MuonIO, a single Muon-style update for both of these layers. For the language model head , we motivate the use of the operator norm, due to the Lipschitz continuity of the softmax output geometry, while for the embedding table , we draw on the operator norm, based on the one-hot input geometry identified by Bernstein & Newhouse (2025). The identity then puts both matrices in the same vocabulary-oriented geometry: MuonIO applies a single normalized-vector rule, which appears as column normalization for and row normalization for . Empirical evaluations demonstrate the effectiveness of our approach, with MuonIO reducing I/O optimizer state memory by 50% and I/O update FLOPs by 46% compared to Muon for 1B LLaMA pretraining on C4, while also improving validation perplexity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.