acceptodds
Under review as a conference paper at ICLR 2027

FISMO: FIsher-Structured Momentum-Orthogonalized Optimizer

Abstract

Training large-scale neural networks requires solving nonconvex optimization problems, where the choice of optimizer strongly affects both convergence behavior and computational efficiency. While adaptive methods like Adam have long dominated practice, the recently proposed Muon optimizer achieves superior performance through orthogonalized momentum updates that enforce isotropic geometry with uniform singular values. However, this strict isotropy discards potentially valuable curvature information encoded in gradient spectra, motivating optimization methods that balance geometric structure with adaptivity. We introduce FISMO (FIsher-Structured Momentum-Orthogonalized), which generalizes isotropic updates to incorporate anisotropic curvature through Fisher information geometry. By reformulating the optimizer update as a trust-region problem constrained by a Kronecker-factored Fisher metric, FISMO achieves structured preconditioning that adapts to local loss landscape geometry while maintaining computational tractability. We establish convergence guarantees for FISMO in nonconvex settings, proving an deterministic rate and an stochastic rate for the expected gradient norm. Moreover, we theoretically show that FISMO can achieve sharper finite-time contraction than Muon-style isotropic orthogonalization through a pre-asymptotic analysis. Empirical evaluation on image classification and language modeling benchmarks demonstrates that FISMO achieves superior training efficiency and final performance compared to established baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.