SWE-Continuum: Multi-Stage Profiling of Behavioral Continuity in Long-Horizon Agents
Abstract
Coding benchmarks usually score the final repository state, omitting when functionality is gained, lost, or recovered. We introduce SWE-Continuum, a benchmark for multi-stage profiling of behavioral continuity, with 16 tasks and 111 requirements from 16 repositories in five languages. Tasks release requirements in a persistent workspace under fixed DAG-based schedules. A verifier checks all released requirements and baseline behavior after each stage. A construction pipeline aligns instructions, tests, and reference patches, with no-op, oracle, dry-run, and human checks. The Area Under the Progress–Expenditure Curve (AUPEC) summarizes net progress over a shared token, time, or cost budget. Across eight models, 17 of 123 valid trials contain 64 pass-to-fail transitions, yet net progress rises at 27 of the 29 checkpoints containing one. Models with nearly equal final net progress differ 4.6-fold in token expenditure; terminal net progress and token AUPEC reverse 14 of 28 model-pair rankings. These results distinguish terminal capability, behavioral continuity, and budgeted net progress.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.