acceptodds
Under review as a conference paper at ICLR 2027

Winner-Take-All: Where Exponential Entropy Collapse Comes From in Reinforcement Learning for LLMs

Abstract

Reinforcement learning with verifiable rewards (RLVR) can improve accuracy while concentrating probability on fewer correct solutions. We study this process through a finite-mode, clean-verifier mean-field model of group relative policy optimization (GRPO), an algorithm used for RLVR. The correct-mode dynamics increase collision probability and decrease both Shannon and Rényi-2 entropy; a unique initial leader eventually captures all correct mass. We derive exact entropy–mass relations along a scalar trajectory and distinguish this mode-level concentration from token entropy through an entropy decomposition over solution families and token sequences. The distinction matters empirically: in a GRPO training run, correct-family diversity falls sharply early in training and only partially recovers, while token entropy subsequently rises above its initial level. Comparing unmodified GRPO with an anti-collision variant, family-count ratios remain between 4.4 and 5.9 over the tested training cutoffs, whereas token-entropy ratios cross unity. The scalar solution also characterizes when accuracy–entropy curves admit shifted-exponential or logistic approximations: each corresponds to a locally constant slope in a different transformed error coordinate. Fits to training curves across model families and sizes on math and coding tasks support both finite-window descriptions. These results motivate measuring correct-family collision alongside token entropy when assessing solution diversity in RLVR.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.