acceptodds
Under review as a conference paper at ICLR 2027

Senior SWE-bench: Evaluating Coding Agents as Senior Software Engineers

Abstract

Coding agents are expected to perform like senior engineers: they are tasked with building features from underspecified requirements, investigating and fixing bugs from user-reported symptoms, and writing high-quality code that meets broader codebase standards. However, popular benchmarks evaluate agents like “junior” software engineers, with highly specified instructions and a focus on behavioral correctness alone. This gap is a consequence of standard grading approaches: pre-written verifiers require full interface specification while LLM-as-a-judge approaches give up the reliability of executing code. Neither can reliably grade senior-level work, which admits many valid solutions that differ in quality. We close the gap with two mechanisms: a verification agent that writes verifier tests adapted to each solution, and a taste judge that assesses code quality in a grounded manner. Using these techniques, we built Senior SWE-bench, a benchmark of 100 tasks (50 public and 50 held-out to resist contamination) drawn from real senior engineering work. Tasks are “design-and-build” (e.g., new features addressing user stories) or “investigate-and-fix” (e.g., diagnosing and solving complex bugs), and instructions are naturally underspecified (on average, 26% of the median length of those in SWE-Bench Pro). Evaluating 46 frontier agent configurations with these grading mechanisms, we find that the best behavioral-correctness pass@1 is 62.1%, while the best overall pass@1 (which also considers taste) is only 33.3%. More broadly, we show that for frontier agents, correctness and taste rankings often diverge, and that Senior SWE-bench tasks can reveal trends in model behavior.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.