COSWE: Evaluating Interactive Planning in Collaborative Software Engineering
Abstract
Coding agents increasingly handle repository exploration, implementation, and testing, but real development often begins with only a partial description of what the user wants. Early benchmarks assessed code against predefined specifications, while recent work has brought requirement clarification and user feedback into the evaluation. This shift calls for measuring interactive planning ability: the speed and quality with which agents identify and formulate the task’s target technical solution through interaction. The target reflects the configured user’s requirements and preferences, without assuming a universally optimal design. We introduce PlanScore as the primary metric, combining discovery timing, resolution delay, and completion to track progress toward accepted decisions. To support this evaluation, we build COSWE, a benchmark based on a dataset of real repository changes. Each task pairs an underspecified request with private reference requirements and a stateful user simulator, so agents must clarify choices, provide behavioral evidence, and explain technical trade-offs as they develop a plan. Report fidelity complements the interaction-based score by assessing whether final reports preserve the reference requirements. Together, the dataset, interaction protocol, and evaluation framework make both the process and outcome of planning observable. Experiments reveal differences in agents’ discovery and resolution performance and show that report omissions consistently exceed contradictions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.