Understanding Encoding in Continuous Diffusion Language Models
Abstract
Continuous language models encode discrete tokens as continuous vectors, but it remains unclear whether encoding design should follow the same principles as in discrete models. In this work, we first show theoretically that encoding has an additional target-defining role in continuous models. Controlled experiments further show how this role leads encoding to affect the two model families in qualitatively different ways. Building on this distinction, we systematically study encoding design and identify geometry, dimension, and encoding–model compatibility as three key considerations. Guided by these insights, we progressively refine a continuous language model's encoding and neural parameterization, achieving a GenPPL of 15.8 on OpenWebText at approximately 130M parameters and outperforming representative state-of-the-art diffusion language models of comparable scale. Our findings establish encoding as a distinct design problem in continuous language modeling and demonstrate the value of systematic encoding analysis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.