SWE-Sweep: Can Agents Autonomously Find and Fix Bugs?
Abstract
Most existing software engineering benchmarks evaluate coding agents on concrete, well-specified tasks, commonly by providing a codebase together with a user-reported issue to resolve. However, as users delegate increasingly broad outcomes to coding agents, the natural next step is for agents to determine not only how to perform useful work, but also what useful work needs to be done. An agent entrusted with a repository should be able to decide what is broken, which problems matter, and how to solve them before they are reported. We introduce SWE-sweep, a benchmark for this open-ended setting. Given a codebase containing many concurrent bugs and no information about their nature or location, an agent must autonomously discover and fix as many bugs as possible. For each repository, we collect real GitHub issue–pull request pairs, then identify a single commit where the maximum number of bugs are present at the same time. Each repair is evaluated against hidden tests from the corresponding pull request, along with the existing test suite to check for regressions. SWE-sweep encompasses 4,068 bugs across 22 programming languages and 100 repositories, including NumPy (Python), OpenCV (C++), and Lean 4 (Lean). We evaluate 7 LMs and find that the best agent resolves only 4.7% of all bugs, with the challenge generally lying in discovering bugs rather than repairing them. SWE-sweep exposes a substantial gap between resolving reported issues and autonomously maintaining software.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.