When Muon Optimizer Meets Adversarial Training: A Theoretical and Empirical Study
Abstract
Adversarial training (AT) remains one of the most reliable empirical defenses against adversarial attacks. Its robustness critically depends on how the underlying min-max objective is optimized. In practice, Stochastic Gradient Descent (SGD) optimizer remains the default optimization choice for AT, whereas adaptive optimizers often improve standard training but may yield inferior robustness. Recently, the Muon optimizer, which orthogonalizes matrix-valued updates via an approximate polar decomposition, has achieved notable success in large-scale training, motivating the investigation of its behavior beyond standard training settings. Robustness claims in AT can be highly sensitive to the optimizer used in the outer minimization, yet optimizer choice is often treated as an implementation detail. A defense may appear robust under one training dynamic but degrade under heterogeneous or adaptive threat models. This raises a question: how does orthogonalized outer optimization alter the robustness and training dynamics of AT, and in which regimes is it beneficial? Focusing on this problem, this paper revisits the choice of optimizer as a security-relevant component of robustness. Theoretically, we show that Muon's polar-factor-based update, together with weight decay, induces a finite worst-case bound on matrix-weight trajectories without explicitly projecting the learned weights. Empirically, across five architectures and multiple threat models on CIFAR-scale settings, Muon often outperforms AdamW and remains competitive with SGD, while SGD remains stronger in some ViT and ImageNet settings. Our results show that outer-optimizer choice can materially change the observed robustness profile under the evaluated AT configurations. These results delineate the regimes in which orthogonalized updates are beneficial. Our findings therefore motivate treating the optimizer as a robustness-relevant design and evaluation variable rather than as an implementation detail.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.