acceptodds
Under review as a conference paper at ICLR 2027

The Answer Is Not the State: Trajectory-Level Corrigibility in Diffusion Language Models

Abstract

Diffusion language models commit tokens across a shared response canvas, making unfinished outputs appear globally revisable. We show that the answer is not the full decision state. Remasking it can leave visible rationale that preserves trajectory history and reconstructs a prior decision despite current evidence. Matched answer and path resets isolate the retained rationale state. Transplanting a coherent rationale between paired trajectories changes the rebuilt decision, whereas shuffling the same tokens weakens the effect. On source-disjoint multi-hop questions, path reset makes decisions in Dream, LLaDA, and iLLaDA responsive to evidence identity, including controlled corrupt evidence. Answer reset retains path dependence, and the absent condition rules out a generic regeneration benefit. In the two-model natural development analysis, direct leave-one-out attribution improves repair prioritization, but even the best partial reset remains far below path reset. Final-answer regeneration alone does not reveal whether a masked-dLLM trajectory remains revisable.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.