acceptodds
Under review as a conference paper at ICLR 2027

Cross-Rollout Bellman Closure for Long-Horizon Agentic Reinforcement Learning

Abstract

Group-based reinforcement learning such as GRPO has become a standard recipe for post-training LLM agents, replacing a learned critic with relative comparison among rollouts sampled for each task. In long-horizon settings, these rollouts revisit shared anchor states, offering cross-rollout evidence for step-level credit. Ideally, step-level credit should incorporate evidence beyond the realized suffixes observed at an anchor while aggregating alternative continuations according to their empirical frequencies. Existing estimators capture only one of these properties: visit-local averaging pools realized suffix returns at shared anchors and respects observed frequencies, but does not recursively propagate evidence across rollouts, whereas shortest-path estimators have global reach but allow a rarely observed route to dominate an anchor's value. Instead, we introduce a Cross-Rollout Bellman Closure (CRBC) method, which merges each rollout group into a finite empirical process with absorbing success and failure boundaries and evaluates its behavior-policy Bellman fixed point with one linear solve. This fixed point uses the same empirical action and transition frequencies to propagate evidence through shared anchors and aggregate alternative continuations. Backing up the resulting state values through observed transitions yields action values, whose gain over the corresponding state value provides step-level credit. A corresponding finite-depth family recovers visit-local return averaging at zero depth and converges to the exact closure as depth increases. The normalized closure credit is combined with the trajectory-level group advantage for policy optimization, without additional environment rollouts. Across ALFWorld, WebShop, and Sokoban benchmarks with multiple model scales, CRBC consistently improves final performance and learning efficiency. For example, CRBC outperforms the strongest baseline by 5.59 percentage points on ALFWorld with Qwen2.5-1.5B-Instruct, achieving a new state-of-the-art performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.