Hindsight Is 20/20: Adaptive Credit Assignment via Realized Futures for Long-Horizon LLM Agents
Abstract
Group-based reinforcement learning (RL) has emerged as an effective critic-free paradigm for training large language model (LLMs) and has recently been extended to long-horizon agentic tasks. However, sparse outcome rewards provide only weak supervision for intermediate decisions, making it difficult to determine which actions are beneficial or detrimental along successful and failed trajectories. Existing step-level group-based methods alleviate this issue through finer-grained intra-group comparisons, but their reliability depends critically on whether grouped actions are genuinely comparable and whether each group provides sufficient support for estimating expected returns. To address these limitations, we propose Hindsight-Enabled Retrospective Policy Optimization (HeRPO), a critic-free framework that elevates realized futures from return signals to posterior context for credit assignment. HeRPO extracts hindsight information from realized futures to construct hindsight-pruned grouping strategy and adaptively adjusts the balancing coefficient in generalized advantage estimation according to the effective sample support of each group. Moreover, for each executed action, HeRPO evaluates the policy both with and without privileged hindsight and converts the resulting shift in action preference into a bounded weight applied to its advantage. By jointly improving the quality of step-level credit signals and controlling the magnitude of policy updates, HeRPO enables more precise optimization for long-horizon agentic tasks. Extensive experiments on the agentic benchmarks demonstrate that HeRPO consistently outperforms other competing agentic reinforcement learning methods. Codes will be available on the GitHub.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.