Granularity and Capacity Matching Policy Optimization for Efficient Agentic RL
Abstract
Training large language model (LLM) agents on long-horizon tasks requires identifying which intermediate decisions contribute to the final outcome. This is challenging when rewards are sparse and delayed. Existing methods trade off the granularity of credit assignment against the cost of value estimation. Token-level PPO provides fine-grained learning signals but incurs substantial critic training costs, whereas outcome-based GRPO assigns the same advantage to every token in a trajectory. We identify a granularity-capacity coupling, suggesting that value-estimation capacity should be matched to the temporal resolution of credit assignment. Based on this insight, we propose Granularity-Capacity Matched Proximal Policy Optimization (GCM-PPO), which jointly aligns credit-assignment granularity, critic capacity, and execution cost. GCM-PPO estimates values at agent-turn boundaries, providing distinct learning signals for successive decisions with substantially fewer value predictions. It employs an adaptive linear-plus-residual critic that reuses detached LLM representations, allowing lightweight CPU-side critic optimization to overlap with GPU-side actor optimization and rollout generation. By adapting critic capacity to actor scale and the available execution overlap, GCM-PPO reduces value-learning overhead while preserving turn-level credit assignment. Experiments across multiple benchmarks and model scales demonstrate that GCM-PPO matches or surpasses strong baselines while substantially improving training efficiency. Our anonymized code repository is available at https://anonymous.4open.science/r/GCM-PPO-C78D.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.