PPO+: Effective and Efficient Policy Optimization with Self-Valuing Transformers
Abstract
Reinforcement learning is central to improving language model reasoning, yet conventional PPO often trails GRPO-style methods and incurs extra cost from a separate critic. At the same time, PPO provides explicit token-level values that can guide learning from sparse rewards, leaving its potential underexplored. We introduce **PPO+**, which improves how value signals guide policy learning, integrates value estimation into the policy itself, and reuses these signals for prompt selection. PPO+ uses both token-level value estimates and within-prompt response comparisons to construct advantages for a single PPO objective. Motivated by our finding that policy representations already contain useful value information, we introduce a *Self-Valuing Transformer (SVT)* that reuses the policy backbone with a small detached value branch. We further introduce *Value-Guided Selection (VGS)*, which combines historical outcomes with early value estimates to prioritize prompts likely to yield mixed groups. Across five mathematical reasoning benchmarks, PPO+ improves over DAPO by 8.09 points on Qwen3-4B while using slightly less overall training time than the GRPO-style baseline. These results suggest that value estimation can become an internal policy capability, enabling stronger and more efficient reinforcement learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.