acceptodds
Under review as a conference paper at ICLR 2027

RODE: A Radial-Orthogonal Decoupled Engine for Optimization

Abstract

Matrix-aware optimizers such as Muon exploit the structure of weight matrices to improve neural network training. However, their additive updates couple changes in weight norm and direction: a directional step can also increase the weight norm, and a larger norm reduces the angular effect of subsequent steps of the same size. This coupling makes it difficult to control weight scale and directional learning independently. We introduce RODE, a Radial-Orthogonal Decoupled Engine that separates each weight matrix into a scalar norm and a unit-norm direction. This separation is nontrivial for Muon-style optimization: as the direction changes, accumulated momentum may no longer lie in the current tangent space, while Newton–Schulz conditioning does not preserve tangency. RODE controls the norm with a scalar radial update and optimizes the direction through tangent-projected Newton–Schulz conditioning and a spherical update, using separate learning rates for the two components. This design allows directional learning without unintended norm growth from additive stepping. Compared with both evaluated Muon variants, RODE achieves lower mean validation loss in GPT-2 and Qwen2-style 114M language modeling and higher mean accuracy on CIFAR-100 and ImageNet-1K using learning rates transferred from GPT-2. RODE also achieves lower validation loss than Muon RMS in Qwen2-style 1.5B pretraining from scratch, and the highest mean accuracy on all four downstream tasks after Qwen3.5-9B fine-tuning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.