Uncertainty Shaping for Risk-Aware Token Fixation in Discrete Diffusion Language Models
Abstract
Discrete diffusion language models (dLLMs) enable parallel generation, but their quality and efficiency depend on which masked tokens are committed at each denoising step: fixing uncertain predictions too early can propagate errors. We study token fixation from a risk-aware perspective and derive a local robustness criterion favoring low-entropy-first fixation under bounded model mismatch and propagation risk. Rather than merely using uncertainty as a decoding score, our framework jointly shapes it during training and exploits it at inference. Differential corruption and structure-aware loss reweighting over a deterministic structural partition encourage structure-associated uncertainty separation, while a dynamic entropy-gated decoder prioritizes low-entropy predictions with per-step fixation constraints. The framework requires no architectural modification or additional model forward pass within a denoising step. On GSM8K, entropy-gated decoding alone yields a **6.1%** relative EM improvement with a **5.64×** speedup over the original LLaDA decoder; the full framework achieves a **23.5%** relative EM improvement and a **9.28×** speedup. Matched controls distinguish entropy ranking from maximum-probability confidence and meaningful structural partitions from length-matched random spans, while cross-task evaluations show improvements beyond GSM8K. Code is anonymously available at [https://anonymous.4open.science/r/EATO-LLM-4DDD/](https://anonymous.4open.science/r/EATO-LLM-4DDD/).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.