acceptodds
Under review as a conference paper at ICLR 2027

Uncommitted Drafts Track Words, Not Knowledge: Evasion and Representation Monitoring in Diffusion LMs

Abstract

A masked diffusion language model holds, at every denoising step, a distribution over its entire uncommitted answer—a whole-canvas draft that recent work reads as an intrinsic monitorability affordance with no autoregressive counterpart. We ask what a monitor of that draft can see, and whether it survives training against it. We build the first secret-keeping, sandbagging, and belief model organisms in diffusion LMs, on Qwen3-0.6B-diffusion and LLaDA-8B. The draft tracks the text a model is writing—including words it will suppress—not the knowledge that shapes that text. Where the secret is the text's topic, a reader of masked-position mass names a word the model emits in under 12% of runs at 0.72–0.92 (chance 0.05), and names it by denoising step 2, while a black-box reader of the committed text needs 8 to 24 steps. Where text and knowledge diverge, the draft fails with no adversary at all: a sandbagger's draft holds the wrong answer it will emit (0.22–0.52) and the correct one only at chance, and a model that acts on an unstated belief in 70% of its answers, never writing the word, trips a word-level draft monitor in just 14% of them at a 5% false-alarm rate. Against a decoy arm given identical training, a word penalty closes the secret's draft signal by 0.52–0.71 while the emitted hints stay informative, and halves the belief organism's word monitor without changing its behaviour. The knowledge is not lost: it stays legible in late-layer representations—where penalty and decoy are indistinguishable at every layer—and deleting a late-layer direction removes the behaviour (0.97 → 0.12, against 0.92 for a random direction). And the concept lives upstream of the token: in a companion organism that states the belief, blocking every token of the word at decoding—identical weights, identical seeds—leaves the dog-specific ideas untouched (0.667 against 0.684, difference -0.02 [-0.14, +0.11]) against floors of 0.009 and 0.000. Draft monitors watch words; the knowledge is in the representation. Monitor concepts, not words.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.