acceptodds
Under review as a conference paper at ICLR 2027

Structure-Guided Policy Optimization for Long-Horizon Language Agents

Abstract

Large language models (LLMs) have shown promise as sequential decision-makers, yet they still struggle in long-horizon tasks where sparse terminal rewards provide limited guidance for intermediate actions. Existing optimization methods mainly densify supervision through trajectory-level aggregation or local scalar proxies, often treating multi-turn interactions as flat sequences. As a result, they lack an explicit coordinate system for measuring task progress, making it difficult to distinguish redundant local exploration from structural transitions that genuinely advance the task. We propose Structure-Guided Policy Optimization (SGPO), a framework that grounds policy learning in the empirical topology of agent-environment interactions. SGPO aggregates self-generated trajectories into a frequency-weighted state transition graph and partitions it into topological regions that capture latent task structure. Anchored by successful outcomes, this recovered structure induces hierarchical guidance signals: an inter-regional signal that encourages progression across structural bottlenecks toward the target region, and an intra-regional signal that guides efficient local navigation toward preferred exits. These normalized structural signals are integrated at the advantage level with group-based outcome supervision, providing dense stepwise feedback while preserving task success as the primary optimization objective. Extensive experiments on ALFWorld, WebShop and Sokoban show that SGPO improves performance and training stability across diverse LLM scales. The source code is available in the supplementary.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.