acceptodds
Under review as a conference paper at ICLR 2027

SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle

Abstract

As autonomous code agents move toward end-to-end software development, evaluating their practical autonomy becomes critical. Current benchmarks hide friction by testing agents in pre-configured environments, and their static evaluation pipelines frequently fail when parsing fully autonomous trajectories. SWE-Cycle addresses these limitations with 489 rigorously filtered instances. It evaluates agents across three isolated tasks, including environment reconstruction, code implementation, and verification test generation, as well as an end-to-end FullCycle task that integrates all three. In FullCycle, one agent starts from a bare repository and carries its own environment, code, tests, and session state across the issue-resolution process. To reliably assess these complex execution paths, we developed SWE-Judge. By combining static code review with dynamic testing, this execution-capable evaluator reliably assesses functional correctness and corrects systematic misjudgments introduced by rigid script-based evaluation pipelines. We evaluate code agents powered by seven state-of-the-art LLMs across these four tasks. The results reveal a sharp drop in solve rates when transitioning from isolated tasks to FullCycle execution, exposing critical bottlenecks in handling cross-phase dependencies and maintaining code quality. Together, SWE-Cycle and SWE-Judge provide a comprehensive framework for accurately measuring the end-to-end capabilities of autonomous software agents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.