Register Tokens for Bounded-State Multi-Block Reasoning in Diffusion Language Models
Abstract
Diffusion language models (dLLMs) promise lower generation latency by denoising many tokens in parallel rather than generating them one at a time. Extending generation, however, typically involves retaining earlier text in context, increasing the cost of each decoding step as the trace grows. We study whether a dLLM can learn to compact its generated state into a fixed number of tokens, allowing it to continue reasoning without increasing its active context length. We implement this state as register tokens: dedicated positions whose continuous hidden states the model learns to write after each block of generated text and read while denoising the next. Specifically, we post-train the LLaDA-8B and Dream-7B models to decode a block of text, clear it while preserving the registers, and continue decoding from the prompt and saved register state. Registers outperform discrete text baselines on all six math and code benchmarks for both models, by up to 8.5 percentage points on GSM8K (a 21% relative gain) and 19.5 on MBPP (91%), and they lead a continuous baseline in 10 of 12 comparisons. We conduct a scaling study showing that increasing the number of registers can further improve accuracy. Registers are also compatible with reinforcement learning: after RL fine-tuning, they outperform discrete text baselines by 2.6 and 8.1 reward points on two multi-chunk reasoning tasks. On code, registers offer a better accuracy–efficiency tradeoff than discrete text baselines, achieving higher accuracy at comparable or lower average decoding time.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.