Learning a Reusable Token-Commitment Policy for Diffusion Language Models
Abstract
Diffusion language models refine many token positions in parallel, but efficient decoding requires deciding when to commit proposed tokens. We ask whether token commitment can be learned as a reusable policy over diffusion traces, rather than calibrated separately for each decoding configuration. We introduce TraceLock, a lightweight controller for a frozen diffusion backbone. An intermediate token proposal is labeled future-stable when it matches the final token of its completed trace; a shared contextual scorer learns from these labels on variable-length trace states. A single TraceLock checkpoint per backbone is deployed across the tested window widths and generation lengths without collecting new traces or retraining. Experiments on mathematical reasoning, question answering, and code generation show an improved quality–step tradeoff over heuristic and learned baselines and greater stability than Learn2PD when the decoding window changes. We use one global acceptance cutoff across tasks and configurations, avoiding per-setting calibration; changing that cutoff provides a controllable quality–step tradeoff. These results support treating token commitment as a reusable learned policy over evolving diffusion traces.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.