DCGPO: Dual-Conditioned Group Policy Optimization for Long-Horizon LLM Agents
Abstract
Group-based reinforcement learning (RL) has become an important paradigm for training long-horizon LLM agents, enabling critic-free advantage estimation through relative comparisons within rollout groups. Recent methods improve credit granularity through step-level grouping, comparing actions taken from recurring environment states. However, state-conditioned grouping answers only which action is better under a matched state; it cannot address the complementary action-conditioned question: why does the same action yield different utility across historical contexts? We call this missing comparison dimension the cross-context incomparability of action utility. We propose DCGPO (Dual-Conditioned Group Policy Optimization), which integrates state-conditioned action comparison with action-conditioned cross-context comparison in a unified advantage estimator. For each action occurrence with cross-trajectory matches in a rollout group, DCGPO constructs a weighted reference return from matching occurrences in other trajectories and defines its action-axis relative advantage as the deviation of the current return from this reference. Because this advantage reflects the relative continuation outcome of the same action under an already-formed context, DCGPO propagates it backward, with distance-based discounting, to the preceding decisions that shaped that context rather than assigning it directly to the current action. It then removes components already explained by episode- and state-level advantages and combines episode-level trajectory quality, state-conditioned action advantage, and residual action-axis credit into a three-axis step-level relative advantage. Experiments on ALFWorld, WebShop, search-augmented question answering, and agentic safety show consistent gains across tasks and model scales, highlighting dual-conditioned comparison as an efficient complement to state-axis credit across diverse long-horizon agent settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.