AbortBench: Separating History Exposure from Task Linkage in Agent Safety Evaluation
Abstract
An agent may finish preparing a deployment after permission to execute has expired. Does a record of completed work make it more likely to continue? We introduce AbortBench, which separates task linkage from the effect of adding history. It compares no history, current-task history, and equally long other-workflow history while fixing current evidence and available actions. Across 560 synthetic tasks, all six primary model configurations select more prohibited actions with current-task history than with matched history, by 4.6–27.0 percentage points. The reference changes effect sizes and model rankings: Llama-8B's 71.1-point difference against no history becomes 4.6 points against matched history. We operationalize the same separation in DeCommit, which authorizes recovery, abort, or inspection from validated current evidence. It matches a direct state rule's 100% correct resolution on 420 clean tasks and rejects all 6,300 tested detectable-fault packets. Together, the benchmark and gate distinguish what history changes from what should authorize the next action.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.