Success Is Not Completion: Goal-Closure Monitoring and Corroborated Contracts for Tool-Using Agents
Abstract
An API can report success even when an agent has not achieved the user’s goal. Repeating a check does not resolve this ambiguity if the check shares the original failure cause. We study post-action goal-closure identification: when does the available evidence justify declaring a task complete? We introduce GoalSeal-CG, a training-free monitor that represents ways the goal may still be open, selects observations that distinguish them from completion, and admits only recovery actions that are safe across the re- maining alternatives. The resulting rule is sound relative to an explicit goal contract and fault model; indistinguishable omitted worlds remain an impossibility boundary. We introduce GoalCommitBench to test safety and useful completion jointly under tool, evidence, contract, and recovery faults. Across 3,456 episodes spanning 1,152 matched conditions from nine model configurations, GoalSeal-CG has zero observed dangerous declara- tions at 70.6% declaration coverage, versus 9.1% danger for prompting. On held-out ToolBench-X workflows it reduces faulty danger from 33.9% to 1.7%. Four of six preregistered criteria pass; the two that do not are clean- execution autonomy gains, where the safety-first gate trades added coverage for zero danger rather than missing recoverable closure. Matched mecha- nism and stateful cross-node studies attribute the safety gain to closure admission and show successful recovery or containment outside a deliber- ately indistinguishable common-mode control. QuorumContract recovers required-clause recall to 100% at zero added danger, but automatic contract exactness remains a bottleneck we isolate rather than mask. A prospective 45-task supportive retest of one targeted gate refinement then reaches 0% faulty danger while improving clean safe closure by 26.7 points and passes six of six pre-specified decision criteria across all three model families. The evidence therefore supports model-relative safety and recovery, not a pro- duction or end-to-end goal-understanding guarantee
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.