acceptodds
Under review as a conference paper at ICLR 2027

CRGPO: Context-guided graph refinement for group-based agent reinforcement learning

Abstract

Group-based reinforcement learning (RL) has improved the performance of large language models (LLMs) on agentic tasks. Despite these gains, it remains hampered by sparse and delayed rewards in long-horizon settings, where feedback often arrives only after dozens of interactions. Recent frameworks have shifted from trajectory-level to step-level training to enable finer-grained policy updates. Yet observation-based credit assignment still aggregates visits with the same local observation, overlooking contextual information from preceding interactions that is relevant to state values. Consequently, visits with different prospects of task success can fall into the same comparison group, distorting the credit assigned to their actions. To address this limitation, we propose CRGPO a context-guided graph refinement method for agentic RL. CRGPO combines separately sampled linear trajectories into a state-transition graph by merging visits with the identical local observations, then refines its nodes through increasingly long context windows. At the first level that distinguishes two contexts, it attributes their estimated value difference to earlier actions at a shared decision branch. Policy updates combine these signals with suffix advantages from the final refined graph and episode-level feedback. Experiments on ALFWorld and WebShop with Qwen2.5-1.5B and Qwen2.5-7B show that CRGPOachieves the highest overall success rates among the challenge benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.