Is a Repaired Prompt-Injection Probe Really Repaired?
Abstract
A prompt-injection probe reads a language model's hidden states and flags a retrieved document that carries an instruction the user did not give. When an application changes the interface, the way the task and the document are laid out in the prompt, such a probe can break. The usual fix is a new threshold or a new linear readout, judged by passing the test on attacks against clean documents again. In this paper, we first show that changing only the interface breaks two released TaskTracker probes in 11 of 12 cells, and that a new readout passes that test again in all 11. Then, we show that passing again can hide two failures that also occur without the shift: a threshold that fires on harmless edits of the document, and a score that misses attacks from another source. We find that adding a kind of text to the repair's data removes its failure on held-out text of that kind: a refit fitted on two attack sources, with harmless how-to steps in both its labels and its threshold, set per kind of text, passes all these tests in 78 of 80 random splits of its data (20 per shifted probe-interface pair). But on a new attack source it fails again: it misses at least 0.57 of its attacks, while the same refit without how-to steps catches them but fires on most how-to steps and quoted injections. So a repaired probe should be tested on harmless text and attacks that were not in its data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.