Better Starts for Value-Based LLM Reasoning: Reference-Return Initialized TBRM
Abstract
Trajectory Bellman Residual Minimization (TBRM) enables value-based reinforcement learning for large language model reasoning by interpreting token logits as action values. However, initializing the model from a reference policy does not automatically calibrate its value estimate: action-independent shifts leave token probabilities unchanged but alter the log-partition used to represent value. This ambiguity can create a mismatch between the initial value and downstream task returns. We give a unified value-anchoring framework that organizes TBRM and the trajectory residual used by ReVal as different choices of a fixed prompt-wise initial value. Within this framework, we propose Reference-Return Initialized TBRM (RI-TBRM), which estimates the reference policy's mean return in an offline annotation pass and reuses the resulting value anchor throughout reinforcement learning. The method applies a fixed offset consistently to action and state values, preserving the initial policy and the trajectory Bellman residual functional without introducing a separate critic. Under reference-policy sampling, exact calibration minimizes the initial residual mean squared error among prompt-wise scalar value initializations and removes the value-mismatch term from the initial expected gradient. We further characterize the estimation error introduced by a finite number of independent calibration rollouts. RI-TBRM achieves the highest six-benchmark average on all three evaluated model backbones.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.