When Success Is a Lossy Record: Matched-Success Continuation Evaluation for Persistent Agents
Abstract
Agent benchmarks reduce a completed episode to a verdict such as PASS, discarding the environment that later tasks inherit. We formalize this compression failure as continuation insufficiency: the excess Bayes risk of predicting future outcomes from the verdict rather than the state under declared continuation conditions. Matched-success evaluation forks equally accepted states into identical frozen continuations and separates transition, preservation, downstream utility, and cost. Across 360 controlled coding pairs, later success differs by 27.50 percentage points; a complementary 240-pair enterprise study shows an 8.75-point shift in the same direction. Among 64 naturally produced successful states, a 2.73-point reference-free separation survives crossed randomization (p = 0.0023). On 144 held-out executions in 18 repositories, retaining continuation information avoids 13.19 points of verdict-only choice regret and recovers the historical-state baseline. Replay and recovery identify history compatibility as the principal omitted variable, while a common-snapshot control isolates downstream divergence in Jsonpickle. Success is therefore a lossy record of an agent's work: our protocol detects the missing information, locates its mechanism, and prices the cost of discarding it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.