A Spectral Perspective on Continuous Diffusion Language Models
Abstract
A spectral Fourier perspective provides a natural lens for understanding visual Diffusion Models. This view inspires a broad range of insights and practical methods. Unlike vision, the text domain does not naturally admit a spectral interpretation due to its discrete nature. However, the rise of Continuous Diffusion Language Models (DLMs), which perform diffusion in the representation space of pretrained text encoders, makes the spectral perspective more natural to investigate in the text domain. In this work, we study continuous DLMs through a Fourier lens. We show that the representation spaces of widely used text models, such as Qwen3.8, Gemma 4, and Nemotron-3, exhibit an approximate power-law spectral structure across sequence length, similar to that observed in images. Notably, this phenomenon persists across different model sizes, architectural families, data domains, training specializations, and instruction-tuning settings. We then analyze the information carried by different frequency bands and uncover a key distinction from vision: high-frequency components in text retain substantially more energy and encode important token-level, positional information, rather than merely fine-grained details. Motivated by these observations, we introduce a frequency-aware diffusion strategy that processes low- and high-frequency components separately across diffusion timesteps, enabling dedicated modeling of high frequencies. Across multiple tasks and representation spaces, this approach yields consistent improvements, with the gains tied to better preservation of positional relationships between tokens.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.