acceptodds
Under review as a conference paper at ICLR 2027

SWE-Quality: A Benchmark for Code Quality in Agentic Software Engineering

Abstract

As LLM agents increasingly automate software engineering tasks, a major bottleneck to deploying their code is **code quality**. Low-quality code is hard to read, hard to edit and expand, and may be less reusable. Repeated edits by agents over the lifetime of a repository may degrade its quality. Despite its importance, agentic benchmarks score correctness almost exclusively. We explore cheap-to-compute static metrics for evaluating agentic code quality. To that end, we introduce **SWE-Quality**, a benchmark built on DeepSWE that computes *quality scores* for agent-produced repository-level patches. SWE-Quality supplies **(i)** quality annotations for DeepSWE v1.1: 72 static metrics spanning a broad array of metric families, under 3 file-scope lenses (216 features), for 3,402 completions from 26 models, and the gold patches of its 33 Python tasks; **(ii)** a scoring protocol built for cross-model fairness, utilising within-task robust standardisation and tie-aware scoring; **(iii)** a validated scoring function, the *SWE-Quality Aligned Index*, a learned combination of those metrics that, on held-out tasks, reproduces 80% of the orderings a blind LLM judge (Claude Fable 5.1) gave to 1,204 within-task pairs of completions from 9 of DeepSWE's best performing models, and ranks those models almost exactly as the judge does (Spearman ); and **(iv)** a baseline leaderboard with ablations. By showing that cheap, deterministic and interpretable metrics reproduce most of an independent judge's quality orderings, SWE-Quality provides a scalable instrument for agentic code quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.