acceptodds
Under review as a conference paper at ICLR 2027

Better Starts for Value-Based LLM Reasoning: Reference-Return Initialized TBRM

Abstract

Trajectory Bellman Residual Minimization (TBRM) enables value-based reinforcement learning for large language model reasoning by interpreting token logits as action values. However, initializing the model from a reference policy does not automatically calibrate its value estimate: action-independent shifts leave token probabilities unchanged but alter the log-partition used to represent value. This ambiguity can create a mismatch between the initial value and downstream task returns. We give a unified value-anchoring framework that organizes TBRM and the trajectory residual used by ReVal as different choices of a fixed prompt-wise initial value. Within this framework, we propose Reference-Return Initialized TBRM (RI-TBRM), which estimates the reference policy's mean return in an offline annotation pass and reuses the resulting value anchor throughout reinforcement learning. The method applies a fixed offset consistently to action and state values, preserving the initial policy and the trajectory Bellman residual functional without introducing a separate critic. Under reference-policy sampling, exact calibration minimizes the initial residual mean squared error among prompt-wise scalar value initializations and removes the value-mismatch term from the initial expected gradient. We further characterize the estimation error introduced by a finite number of independent calibration rollouts. RI-TBRM achieves the highest six-benchmark average on all three evaluated model backbones.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.