acceptodds
Under review as a conference paper at ICLR 2027

Muon for RLVR: What Works, Why, and What It Learns

Abstract

Muon is widely adopted for pretraining large language models and has been extended to supervised fine-tuning, but it is rarely used for reinforcement learning (RL), where reports of its value are mixed. We ask whether Muon can be used for RL with verifiable rewards (RLVR), and if so, how does Muon make a difference. With the right recipe, Muon extends the useful training horizon: it keeps improving after Adam-series optimizers plateau and reaches higher accuracy on both an AdamW-pretrained base (Qwen3-4B-Base, 29.17% versus 24.47% macro-averaged Pass@1 against AdamW) and a Muon-pretrained one (Moonlight-16B-A3B-Instruct, 32.60% versus 31.29% macro-averaged Pass@1 against SignSGD). We identify two critical ingredients. First, the matrix passed to Newton–Schulz should be a single linear operator: orthogonalizing fused QKV projections per head rather than as one stored tensor is what lifts Muon above AdamW on Qwen3-4B-Base, because splitting keeps the cross-head agreement already present in the gradient, which fused orthogonalization whitens away. Second, momentum should be removed, because RLVR offers it no persistent direction to average, Newton–Schulz's scale invariance turns the resulting correlation into extra displacement that AdamW would damp, and on-policy sampling makes that displacement unrecoverable. Finally, Muon, like AdamW, learns off-the-principal directions of the base weights. Muon's endpoint, however, is less compressible and relies on the attention value projections for accuracy where Adam-series optimizers do not.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.