acceptodds
Under review as a conference paper at ICLR 2027

TPPO: Tracking-Value PPO For Language Model Reasoning

Abstract

Proximal Policy Optimization (PPO) is widely used in reinforcement learning with verifiable rewards for language-model reasoning. While value evolution along trajectory characterizes relative token-level credit, yet standard critics often fail to track it under sparse terminal rewards. To model this evolution, we analyze the relation between successive prefix values and show that they admit a Bayesian decomposition with additive updates in log-odds space. We use this additive structure to construct Tracking-Value PPO (\method), a cumulative value parameterization that sums position-wise scalar outputs into value logits. These outputs are learned through terminal-outcome supervision on the reconstructed values. % \method uses the same terminal outcomes as PPO, without additional rollouts or process-level supervision. % Across mathematical reasoning benchmarks, \method consistently outperforms PPO across model scales and variants, improving aggregate scores by 8.5 and 9.2 percentage points on Qwen3-4B-Base and Qwen3-8B-Base, respectively. These gains extend to countdown, science, and coding tasks. Further analysis confirms more accurate value-change prediction and advantage estimation. \method improves aggregate mathematical reasoning scores over PPO by 7.1 and 8.4 points on Qwen3-4B-Base and Qwen3-8B-Base, respectively, while improving out-of-distribution generalization. Gains extend to Qwen3-4B, Qwen2.5-Math-1.5B, Qwen3.5-9B and to countdown, and scientific coding tasks. Further analysis shows more accurate tracking of value evolution and more reliable advantage estimates.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.