acceptodds
Under review as a conference paper at ICLR 2027

Distance-Aware Muon: Adaptive Step Scaling for Normalized Optimization

Abstract

Muon and related normalized optimizers determine a geometry-aware update direction but leave its magnitude to be tuned. Existing stochastic guarantees for star-convex Muon set this magnitude using the unknown distance to a minimizer. We formulate the missing scale as a one-dimensional online-learning problem and introduce Distance-Free Muon, which uses weighted follow-the-regularized-leader to learn the corresponding radius from past stochastic estimates of directional derivatives. For smooth star-convex objectives in general norm geometries, it achieves expected last-iterate suboptimality, matching the known-distance method in its dependence on . The method requires neither the distance to a minimizer nor the noise level and assumes neither a bounded domain nor a bounded trajectory. A complementary predictable stepsize rule gives an bound on the expected gradient norm for general smooth non-convex objectives. Controlled matrix experiments quantify both the finite-horizon cost of learning the radius and its benefit when a fixed radius is transferred across distance scales. For practical additive variants beyond the analyzed updates, separate GPT-124M/WikiText-103 studies show that FTRL rules match tuned Muon, while a current-gradient stepsize-search proxy with a Nesterov-style momentum direction improves an otherwise identical fixed-scale baseline tuned on a disjoint seed.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.