acceptodds
Under review as a conference paper at ICLR 2027

OrScale: Orthogonalized Optimization with Layer-Wise Trust-Ratio Scaling

Abstract

Muon fixes the direction of every matrix-valued update at the polar factor of its momentum, while each layer's step magnitude is addressed only by a static shape correction. We derive a dynamic per-layer scalar by adapting the LARS/LAMB trust-ratio principle to the orthogonalized setting, where the standard denominator candidates—the raw momentum norm or the polar-factor norm—either live in the wrong unit space or carry no update-scale information. The resulting method, OrScale, uses the norm of the parameter-space direction actually applied and anchors each layer's ratio at one via a per-layer calibration, so that the Moonlight recipe (tuned for AdamW, shared with Muon via RMS matching) transfers with no additional sweep; a component ablation confirms each design choice is individually load-bearing. Theoretically, OrScale retains a nuclear-norm convergence rate for any clipped multiplier and achieves a strict layer-adaptive descent gain under two conditions estimable from standard training diagnostics—a bound that predicts the gain should grow with architectural heterogeneity. Experiments confirm the prediction: with every hyperparameter inherited verbatim from the Moonlight recipe, OrScale matches or beats Muon+Moonlight across dense 125M–1.1B FineWeb-Edu pre-training, and on a 16B-A3B mixture-of-experts model—where the logged trust ratios separate cleanly by layer class—the gap widens by an order of magnitude to nats ( relative) at parity wall-clock cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.