acceptodds
Under review as a conference paper at ICLR 2027

SWE-Continuum: Multi-Stage Profiling of Behavioral Continuity in Long-Horizon Agents

Abstract

Coding benchmarks usually score the final repository state, omitting when functionality is gained, lost, or recovered. We introduce SWE-Continuum, a benchmark for multi-stage profiling of behavioral continuity, with 16 tasks and 111 requirements from 16 repositories in five languages. Tasks release requirements in a persistent workspace under fixed DAG-based schedules. A verifier checks all released requirements and baseline behavior after each stage. A construction pipeline aligns instructions, tests, and reference patches, with no-op, oracle, dry-run, and human checks. The Area Under the Progress–Expenditure Curve (AUPEC) summarizes net progress over a shared token, time, or cost budget. Across eight models, 17 of 123 valid trials contain 64 pass-to-fail transitions, yet net progress rises at 27 of the 29 checkpoints containing one. Models with nearly equal final net progress differ 4.6-fold in token expenditure; terminal net progress and token AUPEC reverse 14 of 28 model-pair rankings. These results distinguish terminal capability, behavioral continuity, and budgeted net progress.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.