When Is Self-Repair No Longer Possible? Certifying Irrecoverable Turning Points in Tool-Using Agents
Abstract
A failed repair attempt does not reveal whether a tool-using agent has exhausted every legal recovery path or merely selected a poor action. Existing evaluations often collapse environmental irrecoverability, resource insufficiency, recognition errors, planning errors, and execution errors into one end-to-end score. We introduce an evaluation framework for irrecoverable turning points that specifies the recovery goal, state, action catalogue, and resource budget as an executable contract. A replayable successful path certifies ecoverability, whereas a forward-closed goal-disjoint set certifies irrecoverability. Across 671 checkpoints in finite-state environments, every label agrees with independent search. Contract-sufficiency evaluation verifies 671 full certificates and 945 action-ablation certificates; hiding the remaining budget, capability flags, or domain state creates 9, 42, and 70 ambiguous observation groups. On 24 stateful AgentDojo and AppWorld task clusters, GPT-5.6 Sol, DeepSeek-V4.1-Flash, Qwen3-4B, and Claude Opus5.5 complete 96/96 real-tool repairs. Their exact turning-point results are 24/24, 23/23 among scorable cases, 0/24, and 24/24, respectively. A recognition–planning–execution decomposition further localizes failures concealed by end-to-end outcomes. The framework turns claims that an agent can no longer self-repair into scope-bounded, certificate-backed, and stage-diagnostic evidence for auditable stopping and escalation decisions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.