acceptodds
Under review as a conference paper at ICLR 2027

Not All Tokens Are Equal: Importance-Aware Masking for Discrete Diffusion Language Models

Abstract

Masked discrete diffusion language models (DLMs) corrupt every position with the same probability, so a content word carrying eight bits of surprisal and a function word carrying two are treated identically by the forward process. We introduce ATAT, which makes the corruption depend on which token occupies a position and on when it is corrupted: a continuous importance score preserves informative tokens as context at high noise and concentrates training signal on them at low noise. Two choices make this work. The curriculum is applied as a per-token time warp of the survival probability, expressing the masking probability exponentially through a time-warping exponent on the base survival schedule, rather than as the multiplicative rescaling used by prior non-uniform schemes; for the log-linear schedule and small curriculum bounds we prove this keeps the process a valid absorbing-state chain reaching the fully masked state generation starts from, and we measure rather than assume its masking budget, which departs from the uniform rate by at most 0.014. And the signal comes from the model itself: a masked diffusion model already predicts, at every masked position, a distribution estimating how surprising the underlying token is, so an exponential moving average of the denoiser supplies both representation and supervision, and no autoregressive teacher is needed for supervision, initialisation or inference. In a controlled study holding architecture, data order, optimiser and token budget fixed, importance-aware masking reduces perplexity by 7.4%. That headline does not survive a compute correction: ATAT costs 2.18 times more per training token, and a uniform baseline given the same wall clock leaves 5.2%. We report the contrasts behind it rather than defend it whole: the conditioning input is worth 0.3-0.4 PPL, and placing the curriculum in the exponent rather than multiplicatively is worth 1.0 PPL at matched cost and masking budget, the one comparison the compute correction does not touch.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.