Binary Diffusion Language Model
Abstract
Masked diffusion replaces corrupted tokens with the same mask symbol, removing all local identity information. Instead of representing a token with one-hot vector, one can also represent them with binary vectors. A binary code makes corruption gradual: partial evidence about token identity survives at every noise level. We introduce BinaryDLM, a diffusion language model that flips bits in token codes learned through language-modeling objective, so that code similarity reflects lexical similarity. During training, we sample noise independently across positions, mixing near-clean and highly corrupted states so that informative neighbors appear alongside difficult prediction targets. During generation, high-confidence predictions are temporarily anchored as clean context for their neighbors, and blocks are refined iteratively with early stopping. Across 0.3B–4B models trained on a shared math mixture for up to 100B target tokens, BinaryDLM outperforms both masked diffusion (+0.7–10.2 points) and autoregressive baselines (+2.4–8.5 points) on GSM8K accuracy. Adapting pretrained Qwen3-8B to binary diffusion yields competitive performance with state-of-the-art masked diffusion models of the same scale. These results suggest that the corruption representation is a key design axis for diffusion language models, complementary to the denoising architecture and noise schedule.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.