acceptodds
Under review as a conference paper at ICLR 2027

SafeClawBench: Separating Semantic Failure from Response-Visible Evidence in Tool-Agent Responses

Abstract

Evaluating safety policies requires distinguishing semantic violations from evidence present in model responses. We introduce SafeClawBench, a text-only benchmark of 600 tasks across six attack families, evaluated under four prompt-policy conditions (D0, B2, D3, and D4) on seven models. The benchmark separates semantic rubric failure (CoreFail) from response-visible evidence (RVE): auditor-labeled grounded disclosures or assertions of access, action, or persistence. It evaluates returned text rather than tool execution. Relative to D0, all 21 model–policy comparisons reduce RVE, and 20 are significant after within-model, endpoint-wise Holm correction. D3 and D4 reduce CoreFail in all seven models, whereas B2 does so in six. In the original five-model panel, the reference audit labels 587 of 2,201 CoreFail responses (26.7%) as RVE. Within this panel, each of three auditors assigns more asserted-operation labels than disclosure labels, but their disclosure judgments differ substantially. The lowest observed counts are split between D3 and D4 across models and endpoints, supporting model-dependent policy profiles rather than a universal winner. Together, these findings support evidence-aware policy evaluation that separately reports semantic failures, response-visible evidence channels, and auditor sensitivity, instead of treating them as interchangeable measures of harm.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.