FC-SWE: Failure-Conditioned RL for Long-Horizon Software Engineering Agents
Abstract
Repository-level software engineering (SWE) is a challenging long-horizon setting: agents must reason over extended interactions, use tools across multiple turns, and adapt to observations from stateful environments. Recent work trains SWE agents with reinforcement learning methods such as Group Relative Policy Optimization (GRPO), which independently sample multiple trajectories per issue, test the resulting patches, and compare terminal rewards within a fixed group. However, this training setup does not reuse verifier feedback from failed patches as context for subsequent attempts, even though this feedback contains valuable diagnostic information about what went wrong. Training on recovery trajectories is challenging because the preceding outcome determines whether the next trajectory is generated, while the failed execution determines its conditioning context. We introduce FC-SWE, a failure-conditioned RL framework that incorporates recovery attempts into policy training. After a patch fails verification, FC-SWE restores the repository to its original task state and uses the failed patch and verifier feedback as context for a recovery trajectory. Each attempt is a complete, multi-turn tool-use trajectory with its own verifier reward; a later success leaves the reward assigned to an earlier failed attempt unchanged. FC-SWE adapts GRPO to estimate advantages from the trajectories actually executed for each issue. Failed attempts remain part of training, while unexecuted recovery attempts are excluded from the comparison group. On all 500 SWE-bench Verified tasks under a verifier-assisted protocol, FC-SWE with Qwen3.5-4B and SWE-agent achieves 41.7% Resolved@1 and 52.8% Resolved@2, compared with 38.9% and 48.5% for GRPO. Although trained with at most two attempts per chain, FC-SWE reaches 70.7% Resolved@11 under an eleven-attempt test-time budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.