SWE-QA: Benchmarking Coding Agents for Proactive Runtime Quality Assurance
Abstract
Coding agents are increasingly adopted in software engineering (SWE), demonstrating strong capabilities in resolving explicitly declared bugs in large-scale codebases. However, existing SWE benchmarks typically presuppose high-quality issue reports with detailed information, a premise that rarely holds in practice. Consequently, these benchmarks primarily assess code implementation capabilities, while offering limited evaluation of quality assurance abilities. To address this gap, we introduce SWE-QA, a benchmark for evaluating coding agents on proactive runtime bug discovery, covering 123 quality assurance tasks from 97 real-world repositories across 10 programming languages. Extensive experiments reveal that tasks of discovering latent or multiple bugs triggered by unconventional boundary conditions remain highly challenging for frontier models. We believe SWE-QA provides a complementary perspective on SWE evaluation, and that improved performance on it could help close the gap between autonomous bug discovery and resolution, advancing towards end-to-end software engineering loop for future coding agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.