Do Coding Agents Catch Their Own Bugs? (They Don't)
Abstract
Real software development requires engineers to foresee and fix potential problems with their code that are not explicitly specified in requirements—such as pre- emptively considering vulnerabilities or other bugs from a code change. But current benchmarks often do not measure these future-looking consequences. We introduce **Foresight-SWE**, a benchmark for measuring whether coding agents can anticipate and address unstated problems during software engineering tasks. Foresight-SWE consists of 212 samples spanning 4 languages and 35 popular open-source repositories. Each sample contains two pull requests (PRs), the second of which resolves defects the first one introduced. To measure whether agents exhibit foresight, we define the **foresight resolve rate (FRR)** as the fraction of instances where an agent fully resolves the original task while also avoiding the defect identified by a later PR. We evaluate eleven large language models and find that FRR ranges from only 11.32% to 24.53%, largely failing to address problems introduced by the first PR change. This raises important questions about how models are trained and aligned in coding contexts, as well as how code review tools should be improved for more fail-safe coding. Foresight-SWE shifts evaluations in this direction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.