acceptodds
Under review as a conference paper at ICLR 2027

CAGPO: Critic-Anchored Group Policy Optimization for Multi-Turn LLM Agent Reinforcement Learning

Abstract

Reinforcement learning for LLM agents on long-horizon, multi-turn tasks requires assigning credit not just to whole trajectories but to the specific turns, or even tokens, that actually drive progress. GRPO-style group-relative methods estimate advantage without a critic, but comparing only at the trajectory level discards this fine-grained signal: the policy has no way to tell which actions within a long trajectory mattered. Critic-based methods can recover this signal through value estimation, but jointly training a critic with the policy in agentic settings tends to be unstable and prone to collapse. We propose CAGPO (Critic-Anchored Group PPO), which anchors long-horizon value estimation with a critic while keeping group-relative comparison as the source of fine-grained credit. CAGPO splits the advantage into a macro component, computed via temporal-difference learning along the executed trajectory at either turn or token granularity, and a micro component that runs group-relative comparison among extra candidate responses sampled, but never executed, at selected high-entropy positions. Across several long-horizon agentic benchmarks, CAGPO outperforms both critic-free and hierarchical group-relative baselines, with the largest gains on tasks that require sustained, multi-step completion, surpassing the current state of the art.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.