acceptodds
Under review as a conference paper at ICLR 2027

Tool-Call Safety Evaluation Depends on What Labels Represent

Abstract

A tool-call safety monitor scores proposed calls before execution. Evaluating those scores against eventual attack success can credit or penalize the monitor for later actions. We distinguish whether the pending calls complete the attacker's goal from whether the full run succeeds, applying the benchmark's official success check immediately before and after the calls. These labels disagree on 19.6% of validated public AgentDojo decisions and 26.2% for two newly run agents, mostly because completion occurs later. With predictions fixed, changing the label moves TS-Guard's new-system AUROC from 0.628 to 0.841 and changes held-out monitor selection. ToolSafe's step annotations on a separate benchmark reproduce the sensitivity. On the Qwen agent, at the compared operating points, the locally selected monitor completes 6.1 percentage points more user tasks without attack success than the terminally selected monitor; its attack success is 1.1 points higher, with an interval covering zero. A blocking-disabled arm separates blocking from execution configuration. The event assigned to a score is part of monitor evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.