acceptodds
Under review as a conference paper at ICLR 2027

The Announcement, Not the Attack: What LLM Safety Judges Answer and What Their Hidden States Hold

Abstract

LLM judges are widely used to catch indirect prompt injection, but the answer misses much of what their own hidden states record. To see what the answer and the hidden states each respond to, we separate the two parts of an injection. Besides the task that the user never asked for, an injected tool output usually carries an announcement telling the agent to ignore all previous instructions, and an attacker can simply leave the announcement out. Existing benchmarks mix the two, and their attack runs also differ from the benign ones in length and wording. We therefore rebuild the 62 InjecAgent tasks as pairs of tool outputs that differ in exactly one sentence, the task or a past-tense report of the same event, and add or remove the announcement on its own. We report three findings. First, judges answer the announcement more than the attack. Five open-weight judges separate the announcement from a neutral sentence at 0.967 AUROC or better but the task from its report only at 0.641 to 0.855. Even when the question asks only about actions the user did not request, adding an announcement to a report that carries no task raises the false-positive rate from at most 14.1% to 19.0–90.6%. Second, a linear probe on the hidden states reads the attack itself. Scored on held-out tasks, it separates the task from its report at 0.955 to 0.999 and just as well from a benign imperative of the same length. It also reads whether the user asked for the task, separating authorized from unauthorized tasks at 0.902 to 1.000 against 0.580 to 0.856 for the answer, on records where word, character, style and length classifiers score 0.500. Third, fine-tuning the five judges on announcement-free pairs closes the gap on the task, raising the answer's task detection to 0.987 or better. Judge evaluations should therefore vary the task while holding the announcement fixed.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.