SWE-Bench-XL: Evaluating Coding Agents on Very Large Repositories
Abstract
Coding agents are evaluated on benchmarks of repository-level tasks mined from real pull requests, but current reference benchmarks mostly draw from small codebases. Real industrial and open-source development happens in much larger repositories, where an agent must navigate the code and manage a limited context window. We close this gap with SWE-Bench-XL, a benchmark of 100 repro- ducible tasks mined from four of the largest open-source repositories on GitHub, with a mean per-instance source-code size roughly an order of magnitude larger than reference benchmarks. Across five frontier models evaluated under Open- Hands, pass@1 ranges from 22.5% to 47.6%. Trajectory analysis shows that agents make more codebase searches and visit more directories than in any other benchmark, indicating that SWE-Bench-XL better exercises the capabilities of agents in navigating codebases. We release SWE-Bench-XL as reproducible Har- bor tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.