acceptodds
Under review as a conference paper at ICLR 2027

Shaping the Group via Skill-Probed Policy Optimization for Language Agents

Abstract

Improving language agents on long-horizon tasks hinges on exploration: an agent can only reinforce the behaviors it manages to discover through interaction. In complex environments with delayed rewards, however, breakthrough behaviors stay buried in the low-probability tail of a combinatorial action space, so passive sampling keeps rediscovering the same failures. This bottleneck is especially acute for group relative reinforcement learning such as Group Relative Policy Optimization (GRPO), which forgoes memory-intensive value networks and instead computes advantages from performance differences within a group of rollouts. When the rollouts fail identically, the advantage vanishes and training suffers gradient starvation. To surface the buried behaviors, we propose SPPO, which turns exploration from passive sampling into active group shaping. For each task, an auxiliary generator produces a current skill that approaches the reachable performance ceiling and a lagged skill from a historical snapshot that provides an intermediate coaching floor. Paired with unguided baseline rollouts, these skills prune the search space and shape a trajectory group with a wide reward spread, restoring non-zero, discriminative advantages. The generator itself is trained from environment feedback, with the current and lagged rollouts placed in one group. Crucially, the actor is optimized via a skill-stripped objective with asymmetric credit attribution: positive-advantage trajectories are double-scored to internalize breakthroughs into the autonomous policy, while negative-advantage trajectories are attributed to the skill to protect the standalone policy from probe-induced errors. On a challenging long-horizon agent benchmark, group shaping lowers the fraction of gradient-starved all-fail groups by about points on average over training relative to unguided sampling, and the resulting agent operates completely autonomously at inference time. Our code is available at https://anonymous.4open.science/r/SPPO-shaping-the-group.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.