Reliable and Efficient Agent Loops for Reinforcement Learning on Software Engineering Tasks
Abstract
Software engineering (SWE) agents solve repository-level tasks through tool use, execution, verification, and iterative repair, making SWE an important and objectively verifiable domain for reinforcement learning (RL). Existing benchmarks, agent harnesses, and distributed RL frameworks establish key ingredients, but do not provide a unified infrastructure and training recipe for long-horizon agent loops under resource contention and execution failures. We present SWE-LOOP, which combines a failure-aware SWE agent-loop infrastructure—with bounded admission, class-aware scheduling, image affinity, and asynchronous preparation—with a GRPO recipe using submission-aware rewards, self-verification, group-completion cutoff, masked updates, and sequence-level aggregation. The infrastructure reduces Step-0 wall-clock time by 28.0%, while the full recipe improves SWE-bench Verified and SWE-Bench Pro by 10.8 and 10.6 points over the stable reduced-learning-rate baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.