Measurement Audits of Two Language-Model Attack Detectors
Abstract
An attack detector can reproduce its alerts without measuring the intended behavior. We audit two released language-model detectors by tracing callback inputs, replaying generation prefixes, and testing lexical coverage. DualSentinel's realtime callback repeatedly reads the first token's scores; the lookup appears in every public version we inspected. Crossing the score index with state reset changes verdicts, so an index correction alone does not establish a working detector. On the primary Qwen2.5-14B bank, every tested budget that detects an injection flags legitimate tasks; a separate same-model bank admits detections without legitimate alerts. Qwen-7B traces show budget-invariant legitimate-task alerts. For ICLScan, 248 of 256 keyword subsets preserve benchmark attack alerts because their targets contain redundant wording. A six-family analysis separates target acquisition, elicitation, and recognition. We turn these diagnoses into a practical audit protocol: verify the observed signal, measure legitimate-task costs, and test coverage without redundant cues.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.