SWE-Climb: Benchmarking Coding Agents on Real Work in Million-Line Brownfield Codebases
Abstract
Real-world software engineering rarely starts from an empty repository. Most of it happens inside large, old, heavily modified codebases—brownfield systems accumulated over years of patches, where documentation is stale, configuration is scattered across layers, and behaviors that no document promises are silently depended upon. We present SWE-Climb, a benchmark that operationalizes the three kinds of deliverables engineers actually produce on such systems: explaining them (writing a developer tutorial for a concrete change), running them (authoring executable configs and scripts that drive the repository's real toolchain), and changing them (implementing a reasonable-size feature or fix while preserving old behavior). The current construction snapshot comprises 240 fundamental tasks (120 Tutorial, 120 Config/Script) over 131 real repositories (median 1.6M lines of code and 13.6 years of history, 20 languages), each derived from real upstream development evidence, with de-identified prompts and fully verifiable grading: per-fact rubrics grounded to source paths, offline containerized functional checkers, and fail-to-pass/pass-to-pass behavioral tests. We report three observations: (1) frontier agents leave large headroom—no agent resolves more than 41% of Config tasks, and 29% of them are resolved by no agent in any trial—and two independent grading channels, an LLM-judged rubric and a deterministic binary checker, disagree exactly where explaining and operating a system come apart: the best explainer resolves the fewest Config tasks; (2) running an evolved system is substantially harder than explaining it: every agent's Config resolved rate is 0.36–0.65 of its Tutorial score; (3) failures concentrate on load-bearing facts with no surface signal: for every agent, recall of legacy-invariant facts trails recall of fix-path facts, deep call chains and implicit compatibility paths are the hardest Config mechanisms, repository size barely predicts difficulty, and more tool calls do not buy more resolved tasks. A paired knowledge-injection experiment tests whether supplying exactly these facts closes the end-to-end gap. Alongside the tasks we release the graders, frozen images, version manifests, and full agent trajectories, together with the validity protocol (three-state grader closure, a no-repository probe, adversarial audits, and error-source accounting) that we argue benchmarks of naturally dirty systems require.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.