When Does Momentum Help in Reinforcement Learning with Verifiable Rewards?
Abstract
Momentum is a classical tool for accelerating convex optimization and a standard component of modern large language model training. Yet recent empirical work has questioned the usefulness of first-moment momentum in reinforcement learning with verifiable rewards (RLVR). We study this question through a bandit formulation of RLVR with bias-corrected exponential moving average (EMA) momentum, where the EMA coefficient controls how quickly the weights on past gradients decay. When policy gradients are evaluated exactly, we prove that increasing within a range that prevents overshoot strictly accelerates optimization. When gradients are instead estimated from sampled responses, a positive expected gain can be obscured by sampling fluctuations. We show that as the learning rate decreases, the expected gain with fixed vanishes faster than these fluctuations, whereas increasing to keep within the safe regime preserves a nonvanishing gain that eventually dominates them. Bandit simulations quantitatively validate these predictions. RLVR experiments on Qwen3-4B-Base using SGD with EMA and AdamW support our principle of choosing larger coefficients in practical LLM post-training. On Qwen2.5-32B with AdamW, simply increasing improves mean validation accuracy by nearly percentage points over the conventional choice, demonstrating the principle's effectiveness at a larger model scale.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.