TREEHCA: Hindsight-Guided Exploration And Credit Assignment For Multi-Turn LLM Agents
Abstract
Outcome-supervised group-relative policy optimization (GRPO) gives every token in an LLM agent's trajectory the same advantage, leaving useful and unhelpful decisions undifferentiated. Hindsight credit assignment (HCA) can refine this feedback, but its estimates depend on trajectories containing informative alternatives at consequential histories. We introduce TreeHCA, which couples sampling via hindsight ratios to generate these alternatives with hindsight credit assignment to evaluate their contribution. We use changes in reference-answer likelihood to revisit intermediate histories and select alternative continuations, forming a tree of reasoning and tool interactions. This exploration changes the distribution of sampled continuations. We derive the corresponding inverse-hindsight importance correction, including an exact finite-pool identity, and use its structure in a normalized tree backup that propagates terminal outcomes into turn-level advantages. This common hindsight structure is TreeHCA's central design principle, linking informative exploration to hindsight-guided credit assignment. Across seven question-answering benchmarks and three Qwen backbones, TreeHCA outperforms every baseline in aggregate exact match, validating our paired design for multi-turn LLM training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.