acceptodds
Under review as a conference paper at ICLR 2027

AbortBench: Separating History Exposure from Task Linkage in Agent Safety Evaluation

Abstract

An agent may finish preparing a deployment after permission to execute has expired. Does a record of completed work make it more likely to continue? We introduce AbortBench, which separates task linkage from the effect of adding history. It compares no history, current-task history, and equally long other-workflow history while fixing current evidence and available actions. Across 560 synthetic tasks, all six primary model configurations select more prohibited actions with current-task history than with matched history, by 4.6–27.0 percentage points. The reference changes effect sizes and model rankings: Llama-8B's 71.1-point difference against no history becomes 4.6 points against matched history. We operationalize the same separation in DeCommit, which authorizes recovery, abort, or inspection from validated current evidence. It matches a direct state rule's 100% correct resolution on 420 clean tasks and rejects all 6,300 tested detectable-fault packets. Together, the benchmark and gate distinguish what history changes from what should authorize the next action.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.