Action Context Grouping and Echo Penalization for Agent Policy Optimization
Abstract
Stepwise group relative reinforcement learning has emerged as a computationally efficient paradigm for training Large Language Model (LLM) agents in interactive environments with multiple turns. However, relying solely on environmental observations for state grouping introduces severe credit assignment biases. In this paper, we identify two fundamental optimization bottlenecks in this paradigm centered on observations: (1) spatial degeneracy, where ambiguous observations confound structurally different states, and (2) temporal blindness, where discounted future returns do not explicitly quantify historical redundancy associated with repetitive behavioral loops (the “Echo Trap”). To overcome these limitations without triggering the sample sparsity inherent in dense history matching, we propose Action Echo Policy Optimization (AEPO). At the step level, AEPO introduces Action Conditioned Spatial Disambiguation (ACSD), constructing a soft weighted baseline via action text similarities to precisely isolate spatial noise. At the trajectory level, AEPO incorporates Difficulty Aware Echo Regularization (DAER), a dynamically calibrated penalty that squeezes out trajectory redundancy while preserving necessary exploration. Under matched training and evaluation settings, AEPO improves overall success rates over GRPO, GiGPO, RAGEN (StarPO-S), and ProxMO on ALFWorld and WebShop. With a 1.5B model, AEPO achieves a 93.0% ALFWorld success rate and adds 5.60% to advantage-computation time relative to GiGPO. Additional Search-R1 experiments extend the evaluation to multi-turn retrieval with free-form queries.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.