acceptodds
Under review as a conference paper at ICLR 2027

SWE-Plan: Benchmarking Coding Agents Before They Code

Abstract

Before implementation, current coding agents typically clarify requirements, ask for critical decisions, and form a plan, a process that mainstream products have formalized as Plan Mode. However, existing benchmarks largely entangle this pre-implementation phase with code execution and overlook requirement clarification, leaving this phase unanalyzed and the effects of models and harnesses unattributed. To bridge this gap, we introduce SWE-Plan, the first benchmark that isolates the complete pre-implementation phase and evaluates it directly. Grounded in requirements engineering, SWE-Plan comprises 392 repository-level tasks with 929 withheld critical decisions and 4,937 atomic planning obligations, assessing whether agents recover these decisions and produce complete, repository-consistent plans, with metric reliability validated by human judgment and downstream execution. Evaluating 8 models under 3 mainstream harnesses, we demonstrate that (I) clarification and planning are correlated within a system at r = 0.62 yet not interchangeable; (II) failures in both dimensions concentrate on requirements and verification obligations that are never explicitly stated; and (III) planning rankings correlate at ρ = 0.63 across harnesses, while a 5.5 pp change in ask rate accompanies only a 0.8 pp change in question precision, indicating that the base model mainly determines planning and questioning ability, while the harness mainly shifts the willingness to ask. SWE-Plan provides a diagnostic and systematic evaluation foundation for optimizing coding agents and dedicated planners.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.