CoDev-Bench: Can Coding Agents Handle Naturally Concurrent Development?
Abstract
Current coding benchmarks usually assume a single developer working on a standalone issue. However, real-world software development is inherently concurrent and collaborative, with multiple developers authoring pull requests (PRs) within the same codebase. To bridge this gap, we introduce CoDev-Bench, a novel benchmark designed to evaluate coding agents on concurrent multi-task resolution. CoDev-Bench involves 206 concurrent PR-pair tasks drawn from 137 production-scale open-source repositories across six languages. Each task consists of two independently authored PRs that branch from a shared base commit and are subsequently integrated into the main branch. The benchmark reflects the complexity of real concurrent development, with an average of 17.1 non-test files and 507 lines of code modified per pair, and 45.1% of pairs exhibiting file-level overlap. As reference baselines, we evaluate a single agent that handles both PRs, as well as two multi-agent designs: assigning each PR to a separate agent and merging their patches afterwards, and letting a main agent delegate sub-tasks to sub-agents while it works. Even the best-performing configuration resolves only about one-fifth of tasks on average over three trials, compared with over 80% for state-of-the-art systems on SWE-Bench Verified, showing that naturally concurrent PR-composition tasks remain a largely unsaturated challenge for current coding agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.