CAGPO: Critic-Anchored Group Policy Optimization for Multi-Turn LLM Agent Reinforcement Learning
Abstract
Reinforcement learning for LLM agents on long-horizon, multi-turn tasks requires assigning credit not just to whole trajectories but to the specific turns, or even tokens, that actually drive progress. GRPO-style group-relative methods estimate advantage without a critic, but comparing only at the trajectory level discards this fine-grained signal: the policy has no way to tell which actions within a long trajectory mattered. Critic-based methods can recover this signal through value estimation, but jointly training a critic with the policy in agentic settings tends to be unstable and prone to collapse. We propose CAGPO (Critic-Anchored Group PPO), which anchors long-horizon value estimation with a critic while keeping group-relative comparison as the source of fine-grained credit. CAGPO splits the advantage into a macro component, computed via temporal-difference learning along the executed trajectory at either turn or token granularity, and a micro component that runs group-relative comparison among extra candidate responses sampled, but never executed, at selected high-entropy positions. Across several long-horizon agentic benchmarks, CAGPO outperforms both critic-free and hierarchical group-relative baselines, with the largest gains on tasks that require sustained, multi-step completion, surpassing the current state of the art.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.