acceptodds
Under review as a conference paper at ICLR 2027

Measurement Audits of Two Language-Model Attack Detectors

Abstract

An attack detector can reproduce its alerts without measuring the intended behavior. We audit two released language-model detectors by tracing callback inputs, replaying generation prefixes, and testing lexical coverage. DualSentinel's realtime callback repeatedly reads the first token's scores; the lookup appears in every public version we inspected. Crossing the score index with state reset changes verdicts, so an index correction alone does not establish a working detector. On the primary Qwen2.5-14B bank, every tested budget that detects an injection flags legitimate tasks; a separate same-model bank admits detections without legitimate alerts. Qwen-7B traces show budget-invariant legitimate-task alerts. For ICLScan, 248 of 256 keyword subsets preserve benchmark attack alerts because their targets contain redundant wording. A six-family analysis separates target acquisition, elicitation, and recognition. We turn these diagnoses into a practical audit protocol: verify the observed signal, measure legitimate-task costs, and test coverage without redundant cues.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.