Layer Collapse in Diffusion Language Models
Abstract
Diffusion language models (DLMs) have recently emerged as competitive alternatives to autoregressive (AR) language models, yet the differences in their activation dynamics remain poorly understood. We characterise these dynamics in LLaDA-8B and identify a striking layer-collapse property: early layers exhibit highly similar, collapsed activation patterns, and a single large super-outlier channel dominates the activations at every token position in layers 1–24 (of 32). Despite its apparent redundancy, we have found that this outlier is critical: pruning it causes the model's outputs to degrade into repetitive random token loops. Paradoxically, besides this outlier, layers in LLaDA contain more redundant representations, where redundancy is particularly pronounced in earlier layers. This pattern is the reverse of what is commonly observed in AR language models, where deeper layers tend to become more redundant due to undertraining, as measured by representation similarity. Our analysis further indicates that layer collapse in DLMs is not driven by undertraining. Rather, it appears to arise from overtraining: a dominant outlier becomes an indispensable carrier of information, while the remaining representations collapse into redundant structure. These observations have strong practical implications, as we verify through controlled pre-training experiments. First, DLMs are surprisingly robust to compression: performance of LLaDA under 3-bit GPTQ quantization drops by only 1.8% on GSM8K, whereas Llama-3.1-8B under the same settings drops by 64.7%. Second, optimal non-uniform sparsity allocation reverses between the two model families: under an average budget of 50% sparsity, allocating more sparsity to early layers in LLaDA yields +8.4% over the reverse strategy, while for Llama the same early-layer-sparse allocation incurs -8.4%. Our findings reveal that the DLM training objective fundamentally reshapes layer dynamics relative to AR models, with direct consequences for how such models should be compressed and deployed. Our code will be made available upon publication.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.