acceptodds
Under review as a conference paper at ICLR 2027

Reading Latent States: Adaptive Halting in Continuous Diffusion Language Models

Abstract

Continuous diffusion language models think in vectors, not words, refining a latent state step by step and decoding text only at the end. Yet each trajectory is run to its final call, and two things are unknown. (1) Whether a latent state can be read at all: it is not text, so there is no candidate to score and no confidence to threshold. (2) Whether the full budget is even necessary. To address both, we split the job between a teacher and a student. Offline, a frozen teacher readout marks where the answer enters and whether the state is good enough to stop. At decision time a student reader sees only the latent prefix and a budget signal and emits HALT or CONTINUE—never text or future states. We apply this to CoLA and ELF latent trajectories. The property under test, trajectory parseability, must hold on three counts: predictability from a prefix, information beyond what the computation clock alone gives, and a usable decision. On WMT14 the policy uses 8.8 of 64 solver calls—86% less computation—and returns BLEU 26.2572 against 26.1466 for full inference. On XSum, 9.5 of 64 calls (85% less) at ROUGE-L 27.9232 against 27.811. Unconditional generation is harder: 24.6 of 32 calls, 23% less, at a small perplexity cost. Trajectories make up their minds early: in CoLA the output identity is in place after the first of four latent blocks for 63.6% of saved trajectories, and WMT14 quality peaks at call 19 of 64. Stopping early is thus less a compromise than a correction, because quality is not monotone in computation. Results are replayed from saved trajectories, not a measured speedup.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.