acceptodds
Under review as a conference paper at ICLR 2027

Broadcast, Reinforce, Repeat: Attention Dynamics Behind Repetition Collapse in Diffusion Language Model Forecasting

Abstract

Repetition is a persistent failure mode of language generation, yet its mechanisms in diffusion language models (DLMs) remain poorly understood because bidirectional attention and iterative unmasking introduce dynamics absent from causal decoding. In this work, we study repetition collapse in diffusion-based time-series forecasting, a setting in which predictions can progressively lose temporal variation and converge to pathological numerical repetitions. Through systematic analysis across diffusion steps and Transformer depth, we identify a recurring attention geometry consisting of shared vertical stripes (anchor attention), shifted-diagonal predecessor feedback, and residual structure. We formalize these observations with an idealized reduced dynamical model. Rather than predicting monotonic collapse at every layer, the model shows that positive cumulative anchor–background gain across depth drives anchor concentration and, under appropriate alignment conditions, promotes cross-position representation homogenization. To test this mechanism interventionally, we construct two complementary training-free perturbations that act through distinct computational pathways: attention-level spectral SVD truncation and history-sparsified classifier-free guidance. Both interventions reduce the measured cumulative gain in the predicted direction and substantially suppress repetition while preserving forecasting accuracy. Together, these results connect observed attention geometry, cumulative collapse dynamics, and intervention-based validation into a closed mechanistic account of repetition collapse in diffusion-based time-series forecasting.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.