Bellman Policy Optimization
Abstract
Reinforcement learning with verifiable rewards (RLVR) has become a central approach to improving the reasoning capabilities of large language models (LLMs). We introduce Bellman Policy Optimization (BPO), a critic-free policy optimization method derived from Policy Mirror Descent (PMD). For autoregressive generation with terminal rewards, we use the Bellman equations to reformulate PMD into a trajectory-level objective that avoids estimating intermediate state values. We prove that the reformulated objective has the same unique optimal solution as the original PMD objective. By approximating the trajectory-level objective, we derive the BPO loss, which replaces the importance-sampling ratio in Group-Relative Policy Optimization (GRPO) with a smoothed ratio of complementary token probabilities. In long chain-of-thought and multi-turn tool-integrated mathematical reasoning tasks, BPO achieves the highest average accuracy over the evaluated benchmarks among all compared methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.