From Optimization to Sampling: MUON-Inspired Langevin Dynamics for Bayesian Learning
Abstract
Matrix-aware optimizer such as MUON exploits the spectral geometry of neural-network weight matrices, but their update directions cannot be inserted directly into Langevin dynamics without generally changing the invariant distribution. We ask instead where MUON's matrix geometry can enter Bayesian Neural Network sampling while retaining a clear probabilistic interpretation. We identify two complementary routes. Proximal-MUON Langevin Dynamics (PMLD) transfers the geometry to the target by introducing a spectral prior, yielding a smooth posterior with explicit spectral localization and robustness guarantees. Annealed PMLD (APMLD) takes the transient route: it uses a finite MUON phase to construct a randomized warm start and then switches to target-preserving Langevin sampling. Experiments on Bayesian classifiers and diffusion models reflect these distinct roles: PMLD improves robustness through spectral posterior structure, while APMLD improves finite-time posterior-predictive performance without altering the eventual sampling target. Together, the two methods illustrate a simple design principle for transferring optimizer geometry into Bayesian sampling: change the target, or change only the transient.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.