Abstract-State Grouping Policy Optimization for CLI Agents
Abstract
Recent methods for step-level credit assignment in agentic reinforcement learning, such as Group-in-Group Policy Optimization (GiGPO), compare actions taken from common states across trajectories. However, states of command-line interface (CLI) agents are represented by terminal transcripts that vary across runs. Run-specific paths, hashes, and logs make these transcripts differ even when they describe similar execution situations. Because common-state comparisons require identical transcripts, these differences make common states rare across trajectories. In this paper, we observe that up to 98% of the groups formed by common states contain only one sampled step, and such groups yield zero step-level advantage. To address this problem, we propose Abstract-State Grouping Policy Optimization (ASGPO), a critic-free algorithm that estimates step-level advantages by grouping steps according to abstract states, characterized by recent execution feedback and interaction progress. Specifically, ASGPO maps each preaction transcript to an abstract state, and then groups steps sharing this abstract state and computes their relative advantages using bounded progress-aware returns. These returns combine the final task reward with execution-feedback scores accumulated from each step onward. Positive and negative progress contributions are capped separately to control their scale. Comprehensive experimental results demonstrate the effectiveness of our proposed ASGPO.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.