State-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning
Abstract
Group-relative reinforcement learning removes the need for a learned critic, but trajectory-level advantages provide coarse credit assignment for long-horizon language-model agents. Stepwise group methods refine this signal by comparing actions from a shared anchor state, yet they still score each action using the Monte Carlo return of its own sampled suffix. For long remaining horizons, these suffix returns can accumulate substantial downstream policy and environment stochasticity and cannot exploit future evidence from other trajectories that revisit the same successor states. We introduce State-Graph Policy Optimization (SGPO), a critic-free step-return estimator that shares such evidence across task-matched rollouts. For each rollout group, SGPO merges observed transitions into an empirical state graph, performs certainty-equivalent policy evaluation on the graph, and uses the resulting values to bootstrap a TD() return along each real trajectory. \method shares value information only through actually visited states, requires no additional environment interactions or model evaluations, and leaves the outer group-relative optimizer unchanged. We prove that the empirical Bellman system is well posed for , characterize its behavior-policy value target under a Markov anchor abstraction, and show that recovers the matched Monte Carlo step advantage. Across two Qwen model scales on ALFWorld and WebShop, SGPO achieves higher mean aggregate performance than the strongest evaluated group-based baselines. On WebShop, shuffling successor connectivity substantially reduces performance, providing evidence that the gain depends on the observed transition structure rather than bootstrapping alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.