REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Normalization
Abstract
Critic-free reinforcement learning reduces the cost of post-training large language models, but its behavior depends strongly on how rewards are converted into advantages. In group-relative methods, normalizing rewards separately for each prompt couples the update scale to statistics estimated from a small set of responses. We introduce REINFORCE++, a critic-free policy optimization framework that instead normalizes advantages across the rollout batch. The basic variant supports one or more responses per prompt and incorporates a token-level KL penalty into Monte Carlo returns. A second variant, REINFORCE++-Baseline, combines a within-prompt reward baseline with global normalization and a separate KL regularizer. Our analysis characterizes finite-sample normalization and its convergence to population standardization under independent sampling. Experiments cover preference-based alignment, mathematical and logical reasoning, and multi-step tool use. With one response per prompt, REINFORCE++ obtains an Arena-Hard score of 46.7, compared with 46.8 for four-response GRPO. In tool-use experiments, REINFORCE++-Baseline achieves an average score of 24.10, compared with 22.58 for GRPO and 21.85 for PPO. These results support global normalization as a practical design choice for critic-free language-model training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.