FIRMO: Robust Matrix-Aware Optimization for Neural Networks
Abstract
The optimization of large language models (LLMs) remains a critical challenge, particularly as model scaling exacerbates sensitivity to algorithmic imprecision and training instability. Recent advances in optimizers improve convergence efficiency through momentum orthogonalization but suffer from two key robustness limitations: dimensional fragility in orthogonalization precision and vulnerability to outlier-induced noise. To address these robustness challenges, we introduce FIRMO, an optimizer that enhances training stability through dual robustness mechanisms. First, we develop a dimension-robust orthogonalization scheme that calibrates the full Newton–Schulz composition by matrix shape, achieving higher fidelity than Muon across the evaluated shapes under the same five-step budget. Second, we introduce an optimization-robust framework grounded in proximal optimization that adaptively suppresses extreme input magnitudes while preserving entries within the threshold and retaining matrix-wise orthogonalization. Experiments on a 1B-parameter LLM demonstrate faster convergence in training steps and lower final loss than Muon in 10B-token runs. After 100B-token pre-training, FIRMO achieves 60.12% average zero-shot accuracy across nine tasks, versus 59.59% for Muon and 59.05% for AdamW. Our work establishes a unified framework for robust and precise matrix-aware optimization by jointly exploiting matrix-shape structure and adaptive input-scale control.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.