acceptodds
Under review as a conference paper at ICLR 2027

Beyond Replay Fidelity : Restoration Completeness in Agent Evaluation

Abstract

Resuming a tool-using agent requires recovering the state on which its next actions depend. A checkpoint can faithfully preserve a record of success while losing the application effect it records. We study restoration completeness as a condition of agent evaluation, using a decision-contract framework that distinguishes local information sufficiency, native-score sensitivity, and adaptive recovery effort. Enumerating all 19 AppWorld development families (57 instances) yields 28 qualified boundaries. A controlled intervention retains the visible checkpoint while omitting application effects; native scores fall in nine families (25 instances), and replaying the omitted statement rescues all 25. A 180-episode study evaluates GPT-5.6-sol and GPT-5.5 separately, each with five fresh pairs per instance on the same nine selected affected instances. Complete restoration succeeds in 45/45 episodes per policy; Partial succeeds in 40/45 and 38/45, respectively. Mean paired increases are 1.82 and 0.87 Python execution requests (conditional 95% bootstrap intervals [1.42, 2.24] and [0.20, 1.53]), or 9.1% and 4.3% of the 20-request budget, computed before rounding. Both policies register completion in all five of their Partial reminder failures. Complementary Cybench witnesses test strict action-ranking reversals. Average added effort recurs across the two evaluated campaigns, while its magnitude and family-level signs vary. Evaluations should declare the restoration boundary and report native reward together with recovery effort under the stated interface and budget.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.