Hierarchical Text Generation with Next-Concept Modeling (NCM)
Abstract
Diffusion language models (DLMs) are appealing for their parallel sampling, which can enable faster generation than autoregressive models; this is increasingly important as language model inference has become a dominant computational workload. In practice, however, DLM quality degrades quickly as the number of sampling steps shrinks, every step processes the full sequence, and learning to generate in every possible order makes training expensive. We introduce , a two-stage, coarse-to-fine framework in which a multiscale residual VQ-VAE encodes text into a few levels of discrete codes and a transformer generates them one level at a time, predicting all codes within a level in parallel. At the M-parameter scale, achieves a better generative perplexity on LM1B than the strongest DLM baseline while using fewer steps and running faster. On OpenWebText, it approaches the generative perplexity of the strongest diffusion baseline while running faster, and on XSum it matches the strongest diffusion baseline in ROUGE-L while sampling faster.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.