Geometric-Steering Muon for Online Large Language Model Fine-Tuning
Abstract
Fine-tuning is a fundamental approach for adapting pretrained large language models (LLMs) to downstream tasks and evolving application requirements. Although Muon has recently attracted increasing attention as a matrix-structured optimizer for LLM training and fine-tuning, existing studies have largely focused on conventional offline settings, leaving its behavior under online adaptation comparatively underexplored. In online LLM fine-tuning, evolving objectives can render accumulated momentum increasingly inconsistent with the current optimization signal, yet Muon's Newton–Schulz (NS) transformation primarily reshapes the spectral structure inherited from momentum without explicitly adapting its update orientation to the current objective. We therefore propose Geometric-Steering Muon (GS-Muon), which introduces Geometric-Steering Newton–Schulz (GS-NS) to explicitly couple Muon's spectral transformation with the current optimization signal. GS-NS decouples orientation adaptation from spectral orthogonalization: a short NS prefix establishes a near-orthogonal geometry, a tangent-inspired steering step redirects the momentum-inherited orientation toward the current optimization signal, and a final NS correction restores the orthogonalization structure. Theoretically, we establish geometric stability properties of GS-NS and derive variation-dependent nonconvex regret guarantees for GS-Muon under both first-order and zeroth-order signal access, explicitly characterizing the effects of momentum tracking and zeroth-order signal estimation. Finally, we establish a unified streaming evaluation protocol across two model scales and eleven NLP datasets, covering single-dataset streams, sequential within-family task shifts, and heterogeneous cross-task streams.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.