SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integration
Abstract
Large language model (LLM)-powered agents have demonstrated strong capabilities in automating software engineering tasks such as bug fixing. However, in the real world, software engineering is an inherently iterative process driven by evolving requirements and long-term feature development. Static, one-shot evaluation paradigms fail to capture this dynamic. To bridge this gap, we propose SWE-CI, a novel repository-level benchmark grounded in the Continuous Integration loop, designed to shift the evaluation paradigm for coding agents from static, short-term functional correctness toward dynamic, long-term regression-aware code evolution. The key insight is simple: A testable aspect of maintainability can be revealed by tracking how functional correctness changes across successive code modifications. The benchmark comprises 100 tasks, each derived from a real-world code repository with a development history spanning an average of 233 days and 71 consecutive commits. SWE-CI requires agents to systematically resolve these tasks through multiple rounds of analysis and coding iterations. Extensive experiments shed light on regression, a problem systematically overlooked by existing benchmarks: an acceptable final deliverable does not imply an acceptable development process.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.