acceptodds
Under review as a conference paper at ICLR 2027

SpecAlign: Test-Driven Specification Refinement for Coding Agents

Abstract

Software specifications, such as issue descriptions, are usually informal and often vague. Nevertheless, users insist that code written or generated based on a specification must align with it. This may require clarifying the specification itself. Prior work attempted to improve code+spec alignment by automatically rewriting the specification in a standardized way. Unfortunately, doing so just shifts the problem: now it is burdensome for users to tell whether the standardized spec aligns with the original. Instead, we propose SPECALIGN, which takes a different approach: it pinpoints parts of the spec to improve and surgically rewrites only those parts. That makes it is easier to compare the original and rewritten spec. To align a spec with generated code, SPECALIGN generates both code and tests, and then optimizes how well they agree. We demonstrate that this leads to better spec+code alignment, and ultimately, better issue resolution rates on SWE-Bench Pro and SWE-rebench. Even without a human, just pinpointing the gap already yields 2.4-3.6% improvement while only touching 35% of words. To understand the remaining gap, we introduce a simulated user into the SPECALIGN loop. Thanks to SPECALIGN pinpointing the gap, a simulated human can get 13.8-34.4% improvement. We quantify how often it requires information missing from the spec and available only in the ground truth. Overall, we hope that this paper contributes to better aligning AI-generated code with human intentions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.