Rethinking Tokenization for Discrete Diffusion Language Models
Abstract
Discrete diffusion models offer a potential speed advantage over autoregressive models by generating multiple tokens per denoising step. However, parallel decoding samples each position independently within a denoising step. This ignores the dependence between tokens generated in the same step. This discrepancy, known as factorization error, grows as more tokens are decoded in parallel. Prior work has sought to reduce this error through sampling schedules or auxiliary models, but these approaches address which tokens are generated together without changing the tokens themselves. We instead investigate tokenization itself as a design choice that shapes this error. Discrete diffusion models commonly inherit tokenizers designed for autoregressive models, despite token boundaries determining which dependencies remain between independently predicted outputs. Lossless tokenizers sharing vocabulary size and per-example encoded lengths achieve the same oracle denoising loss, yet they can incur different factorization errors under parallel sampling. We decompose this error exactly by interaction order, separating schedule-determined weights from tokenizer-determined interactions. This decomposition clarifies when pairwise dependence suffices for comparing tokenizers and when higher-order interactions alter the ranking. We compare BPE, WordPiece, Unigram, and a weighted pointwise mutual information (WPMI) merge criterion motivated by the adjacent pairwise component in this decomposition. Experiments on LM1B and OpenWebText show that tokenizers with similar compression and prediction loss can diverge in generation quality as the number of denoising steps decreases. Our findings suggest that tokenizer selection, commonly inherited from autoregressive language-modeling pipelines, may need to be rethought for discrete diffusion language models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.