acceptodds
Under review as a conference paper at ICLR 2027

WHEN TOKEN-LEVEL DIVERGENCE MISRANKS PROMPT-INJECTION ROBUSTNESS

Abstract

Masked diffusion language models expose a predictive distribution at every output position, which makes the Kullback–Leibler (KL) divergence between a clean run and a prompt-injected run a convenient score for comparing decoders. We ask whether this score can rank decoders by their robustness to prompt injection. We study Dream-7B and LLaDA-8B on templated record-lookup cases and on 48 single-action tool cases in AgentDojo, averaging the KL of each output position at the step where that position is fixed. On the 32 primary lookup cases, Dream with a raised end-of-sequence (EOS) logit has mean KL 0.080 nats and follows the injection in 9 cases, whereas Dream with a penalty on the attacker’s target tokens, an oracle that assumes the target is known, has 1.101 nats and follows it in 0. Most of this KL comes from answers that keep the correct label but drop the sentence around it. A control that varies label length and requested answer format independently orders these two decoders correctly, but across its four Dream decoders every decision-blind score we tested has rank correlation −0.40 with attack counts, including the mean after subtracting the KL caused by a matched benign note, which repairs the focal comparison in both panels. Changed tokens still carry almost all of the KL; which of them matter depends on the decision rule. We show that per-position divergences cannot identify decision preservation without the decision rule, and give a margin condition under which stable decision-relevant tokens preserve an authorized action. A KL score says something about safety only once we know which output tokens decide the action.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.