Demystifying Manifold Constraints in LLM Pre-training
Abstract
The recent success of matrix optimizers (e.g., Muon) suggests that specific normalization of momentum—such as orthogonalization and row-wise normalization—benefits both the stability and acceleration of LLM training. Consequently, several recent studies have suggested that weights should also be normalized, leading to a Riemannian optimization problem. While such constrained training frameworks demonstrate superior performance, the effects of explicitly constraining weights, and their interaction with existing stabilization mechanisms, remain less understood. To bridge this gap, we study manifold constrained training dynamics through activation scales, rotational dynamics, and the update-to-weight ratio. We propose a Riemannian spectral steepest descent optimizer called MACRO, alongside a radius selection principle to serve as our testbed. Our analysis and numerical experiments reveal that RMSNorm and manifold constraints serve overlapping roles, and that weight decay can be completely eliminated when manifold constraints are applied. By controlling the update-to-weight ratio, constrained training significantly alleviates update cancellation, empirically demonstrating that MACRO is robust to low-precision computation and competitive with existing algorithms for standard LLM pre-training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.