Diversity-Preserving RLVR via Information Projection
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR), exemplified by Group Relative Policy Optimization (GRPO), is widely used for post-training of modern generative models. Although GRPO improves reasoning accuracy, it can also over-sharpen model output distributions, suppressing alternative correct responses and reducing solution diversity. We mitigate this issue through information projection, seeking the policy that best preserves the response diversity of the base model at a target expected reward. Divergence from the base policy serves as a proxy for diversity loss. To this end, we propose -Bregman Policy Optimization (BPO), a family of representation-based policy updates derived from information-projection dynamics under -Bregman divergences. In a tabular model, we characterize the resulting projection flow and prove that it minimizes the corresponding divergence from the base policy at every attained reward level. Experiments on the Hyper-Grid task illustrate that BPO attains a reward-maximizing policy while preserving diversity among high-reward actions. Extensive experiments on mathematical reasoning further show that BPO maintains competitive Pass@ while improving Pass@ at large sampling budgets, and achieves the KL-divergence-accuracy Pareto frontier, supporting information projection as a principled framework for mitigating over-sharpening during RLVR.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.