Optimality-Preserving Exploration Enhancement via Correctness-Conditioned Entropy
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) relies on exploration to discover high-quality solutions, yet the policy diversity of large language models (LLMs) often collapses early in training. Entropy regularization can alleviate this issue but is widely observed to reduce accuracy. Our geometric analysis attributes this tension to conflicting objectives: RLVR minimizes the projection distance to the optimal policy space, whereas entropy regularization pulls the policy toward a uniform distribution. Since the uniform distribution is generally not optimal, these two distances cannot be minimized simultaneously, implying an inherent conflict. We therefore propose Correctness-Conditioned Entropy (CCE) regularization, which regularizes the policy's projection onto the optimal policy space rather than the policy itself. Because the projection point and the projection distance are independent and can be optimized simultaneously, CCE regularization is optimality-preserving. We further show that RLVR with CCE is equivalent to variationally minimizing the Chernoff -divergence to the maximum-entropy optimal policy, providing a principled interpretation of CCE. Empirically, integrating CCE into DAPO and MaxRL yields consistent improvements across most mathematical reasoning benchmarks. When added to DAPO, CCE improves average accuracy by 1.3–2.0 points while enabling more targeted exploration, achieving simultaneous gains in exploration and optimality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.