Adaptive Advantage Redistribution for Software-Engineering Agents
Abstract
Reinforcement learning for software-engineering agents uses terminal test outcomes to supervise trajectories of commands and edits. Episode-level advantages provide uniform supervision within each trajectory, limiting step-level credit assignment. We propose adaptive advantage redistribution, which groups steps across trajectories by their pre-action state. Local advantages are then estimated by comparing the terminal returns associated with different actions taken at the same state. When only a few rollouts visit a state, its mean return can be unreliable. Adaptive shrinkage combines this state-group mean with the task-level mean to estimate the comparison baseline. Direct addition of the resulting local advantages can alter the original sum of token advantages by different amounts across trajectories, changing their relative contributions to the policy update. We address this through constrained redistribution, which converts local advantages into step-specific corrections while preserving each trajectory's original advantage sum. The policy objective then clips the episode-level and correction terms separately, retaining a restoring gradient when their signs oppose and the probability ratio exceeds the clipping bounds. With Qwen3-Coder-30B-A3B-Instruct, the method achieves resolve rates of 57.0% on SWE-bench Verified and 26.0% on SWE-bench Pro Public, compared with 50.8% and 21.9%, respectively, for episode-only training. Ablation experiments demonstrate the effectiveness of step-level credit assignment, adaptive shrinkage, constrained redistribution, and separate clipping.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.