RODE: A Radial-Orthogonal Decoupled Engine for Optimization
Abstract
Matrix-aware optimizers such as Muon exploit the structure of weight matrices to improve neural network training. However, their additive updates couple changes in weight norm and direction: a directional step can also increase the weight norm, and a larger norm reduces the angular effect of subsequent steps of the same size. This coupling makes it difficult to control weight scale and directional learning independently. We introduce RODE, a Radial-Orthogonal Decoupled Engine that separates each weight matrix into a scalar norm and a unit-norm direction. This separation is nontrivial for Muon-style optimization: as the direction changes, accumulated momentum may no longer lie in the current tangent space, while Newton–Schulz conditioning does not preserve tangency. RODE controls the norm with a scalar radial update and optimizes the direction through tangent-projected Newton–Schulz conditioning and a spherical update, using separate learning rates for the two components. This design allows directional learning without unintended norm growth from additive stepping. Compared with both evaluated Muon variants, RODE achieves lower mean validation loss in GPT-2 and Qwen2-style 114M language modeling and higher mean accuracy on CIFAR-100 and ImageNet-1K using learning rates transferred from GPT-2. RODE also achieves lower validation loss than Muon RMS in Qwen2-style 1.5B pretraining from scratch, and the highest mean accuracy on all four downstream tasks after Qwen3.5-9B fine-tuning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.