acceptodds
Under review as a conference paper at ICLR 2027

Muon on the Quotient: Gauge-Equivariant Descent for Homogeneous Networks

Abstract

Positively homogeneous networks, including ReLU networks and the SwiGLU blocks of modern transformers, have a rescaling symmetry. That is, scaling up the incoming weights of a hidden unit and scaling down its outgoing weights leaves the function unchanged. Orthogonalized optimizers such as Muon do not respect this symmetry, thereby equivalent parameterizations of the same function receive different updates, and a function-preserving rescaling can lead Muon to a sharper basin with a higher loss. We introduce Gauge-Muon, a family of optimizers that keeps Muon's orthogonalized update but applies it to gradients rescaled by per-unit factors computed from the weights, and then undoes the rescaling. Every member of the family is exactly equivariant under hidden-unit rescaling at every finite step size, and performs exact steepest descent in the quotient geometry that identifies all equivalent parameterizations. We further show how the family's exponent allocates functional motion across hidden units, and derive an invariant capacity bound on generalization. Empirically, Gauge-Muon is unaffected by rescalings that degrade Muon and improves on Muon in image classification and GPT-2 language modeling, with the most consistent gains in mixture-of-experts models with narrow experts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.