Terminal Self-Degradation: Code Agents End Worse Than They Reached
Abstract
Code agents are scored on the patch they finally submit. That score hides a failure we can only see by executing what came before it: in a failed trajectory, an earlier executable state can pass a strict superset of the tests the submitted patch passes. We call this terminal self-degradation (TSD). Its strongest form, solved-then-broke, is a trajectory that reached a state passing the entire suite and then submitted one that does not. To measure it we replay every reconstructable intermediate state of fresh full-benchmark SWE-bench Verified rollouts through the official harness. A scaffold is admitted only if its reconstructed states pass a submission-time parity check, and a bash-only scaffold that fails it is excluded. On our primary cell (two frontier GPT-5.6 models on SWE-Agent), at least 4.9% of all failures are TSD. Nine are solved-then-broke: a harness-certified resolving patch was on disk and something else was submitted, and in seven of the nine the agent had itself evaluated that state before editing it away. It is more pronounced on six open-weight model lines (5.7–14.6% of all failures, every one above the frontier floor), and also holds on a second scaffold and eight non-Python language groups, surviving triple re-execution. Degradation proceeds mostly through PASS_TO_PASS regressions that a stable or rising pass count masks. The agents' own tests never separated the lost state from the submitted one. Most primary-cell losses are clean un-fixes that differ from the final only on the held-out target, so a reference-free selector built on repository tests recovers almost none of them. We establish TSD's existence, generality, and structure, report lower bounds rather than prevalences, and argue that the gap between failing to generate a fix and failing to retain one is a diagnostic axis that final-patch scoring, and any signal derived from it, cannot currently report.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.