acceptodds
Under review as a conference paper at ICLR 2027

MileGPO: Milestone Inference with Local Evidence for Graph-Based Policy Optimization of Long-Horizon LLM Agents

Abstract

Credit assignment is challenging in long-horizon agentic reinforcement learning, where supervision typically comes from discrete final rewards. Existing methods refine trajectory-level signals into step-level credits through anchor-based step grouping or graph-based discriminative advantage estimation, but still overlook nontrivial milestones in the intermediate process. To address this, we propose MileGPO (Milestone Inference with Local Evidence for Graph-Based Policy Optimization), developed through three nested credit designs over grouped on-policy rollouts. First, Milestone Discovery identifies candidate milestones from successful rollouts and recurring traps from failed ones. Second, Reliability-Calibrated Shaping (RCS) assigns each candidate a confidence-based weight: high-confidence milestones receive stronger positive credit, high-confidence traps incur stronger penalties, and uncertain candidates are down-weighted. Finally, Progress-Contrastive Calibration (PCC) determines which candidate milestones deserve stronger credit by checking whether they mark local progress and whether the transition reaching them outperforms other observed choices from the same state. Using only existing grouped on-policy rollouts and rewards, MileGPO provides fine-grained process-level credit without external supervision, auxiliary models, or additional environment interaction. Extensive experiments demonstrate that MileGPO achieves superior performance on both the ALFWorld and WebShop benchmarks, while the marginal performance gap between the in-distribution and out-of-distribution ALFWorld test sets further supports its generalization to the ALFWorld out-of-distribution split. Cumulative ablations support the incremental benefits of RCS and PCC in the evaluated configurations, while rollout diagnostics characterize how outcome-derived signals modify intermediate credit allocation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.