GHGPO: Finite-Budget Graph-Hit Credit Assignment for Multi-turn Agentic Tasks
Abstract
Fine-grained credit assignment is critical for reinforcement learning in multi-turn agents, yet critic-free group-based methods typically provide only coarse trajectory-level supervision. Recent graph-based approaches improve process credit by using structural distance to success states, but structural efficiency alone does not fully characterize the downstream success probability under the current policy and the finite interaction budget. We propose **Graph-Hit Group Policy Optimization (GHGPO)**, which retains both the state-transition topology and the empirical frequencies of repeated transitions in grouped rollouts to construct an empirical probabilistic state-transition graph reflecting policy-dependent transition behaviors. Based on this graph, GHGPO estimates a finite-budget graph-hit probability to characterize policy-dependent success probability within the available interaction budget, and uses it for fine-grained process credit assignment. Across ALFWorld, WebShop, and Sokoban, GHGPO consistently improves aggregate performance over strong reinforcement learning baselines across different model scales. Ablation studies further validate the importance of finite-budget estimation and the complementarity between policy-dependent success probability and structural efficiency, while GHGPO requires no additional environment interactions and introduces less than 0.2% additional per-iteration computation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.