Learning to Assign Credit in Hindsight for Long-Horizon LLM Agents
Abstract
Sparse outcome rewards make it hard to tell which actions actually mattered, and training on them directly risks treating pivotal actions and irrelevant ones alike. Existing work has approached this issue through both prospective credit estimation and hindsight prompting. We take a complementary approach: rather than prompt for hindsight, we explicitly learn it. We ask: Given the outcome and the state, how likely was the action that was taken? Comparing this to current policy likelihood provides a fine-grained attribution of each action to the realized future, letting us amplify or suppress its share of reward accordingly. We train a lightweight parameter-efficient estimator of this hindsight conditional, avoiding repeated prompting and reliance on instruction following. We integrate the resulting action-level attribution into the GRPO policy-gradient objective to construct our method, Learned Hindsight Credit Assignment (LHCA). Through extensive experiments and ablations on ALFWorld, Sokoban and BabyAI, we show LHCA outperforms prior credit-assignment methods, including prompted hindsight and preference-trained process-reward estimators, highlighting the effectiveness of explicitly learning hindsight attribution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.