TMuon: Tangent Momentum improves Muon for LLM Training
Abstract
Geometry-aware optimization has recently advanced large-scale model pretraining by more effectively leveraging gradient structure and spectral information. Muon, as a representative example, uses Newton–Schulz orthogonalization to construct update directions from momentum matrices. However, Muon’s standard momentum signal can contain substantial components geometrically aligned with the current model weights. In this work, we propose TMuon, a simple yet effective optimizer that incorporates tangent-projected momentum as input to Newton–Schulz orthogonalization within the Muon framework. Tangent filtering suppresses locally redundant components in the update signal, yielding a more informative momentum matrix for spectral normalization. We further show that norm-dependent tangent filtering leverages local geometry along the optimization trajectory and improves the input to Newton–Schulz orthogonalization. Across pretraining experiments on LLaMA and Qwen3 models, TMuon consistently improves optimization performance over Muon and recent variants at negligible additional cost. In particular, on Qwen3-1.7B trained with a B-token budget, TMuon improves average benchmark accuracy over Muon by percentage points, suggesting a promising direction for LLM optimizers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.