Mano: Restriking Manifold Optimization for LLM Pretraining
Abstract
While large language models (LLMs) have emerged as a significant advancement in artificial intelligence, optimizer design for LLMs has attracted much attention and costs huge resources. In this line of research, advanced optimizers, such as Adam and Muon, with complex operations are mainstream methods, while conventional manifold optimization methods have been largely overlooked due to their poor performance in large-scale model optimization. In this study, we try to restrike manifold optimization for training LLMs. We design a novel Manifold-Normalized Optimizer (Mano) that trains LLMs on a special rotational Oblique manifold implicitly. Mano's Oblique geometry actually corresponds to a product of neuron-wise hyperspheres that make training LLMs easy even with simple momentum. We also theoretically analyzed Mano in terms of convergence and Fisher information. Experimental results demonstrate that Mano outperforms AdamW and Muon across multi-scale LLaMA and Qwen3 models. Meanwhile, Mano incurs lower computational and memory costs than Muon and Adam.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.