Winner-Take-All: Where Exponential Entropy Collapse Comes From in Reinforcement Learning for LLMs
Abstract
Reinforcement learning with verifiable rewards (RLVR) can improve accuracy while concentrating probability on fewer correct solutions. We study this process through a finite-mode, clean-verifier mean-field model of group relative policy optimization (GRPO), an algorithm used for RLVR. The correct-mode dynamics increase collision probability and decrease both Shannon and Rényi-2 entropy; a unique initial leader eventually captures all correct mass. We derive exact entropy–mass relations along a scalar trajectory and distinguish this mode-level concentration from token entropy through an entropy decomposition over solution families and token sequences. The distinction matters empirically: in a GRPO training run, correct-family diversity falls sharply early in training and only partially recovers, while token entropy subsequently rises above its initial level. Comparing unmodified GRPO with an anti-collision variant, family-count ratios remain between 4.4 and 5.9 over the tested training cutoffs, whereas token-entropy ratios cross unity. The scalar solution also characterizes when accuracy–entropy curves admit shifted-exponential or logistic approximations: each corresponds to a locally constant slope in a different transformed error coordinate. Fits to training curves across model families and sizes on math and coding tasks support both finite-window descriptions. These results motivate measuring correct-family collision alongside token entropy when assessing solution diversity in RLVR.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.