Drifting and Merging: The Dynamics of LLM Degeneration
Abstract
An LLM can continue generating long after its output has stopped making meaningful progress. What determines whether generation continues to develop the text or becomes generic, incoherent, or repetitive? Since Holtzman et al. (2020), much research has focused on decoding and fine-tuning methods to reduce these problems, but how they develop during generation remains poorly understood. We organize these forms of text degeneration through two autoregressive effects: *drifting* moves the distribution of generated text away from the natural-language distribution, while *merging* makes predictions from different histories increasingly similar. To test this framework quantitatively, we focus on persistent repetition, whose onset is relatively easy to identify. We call this onset time the *collapse time*. We define and measure drifting and merging times, and , from prompts alone and use them to predict collapse time. Across 16 models, five corpora, and five attention temperatures, a regression using these two timescales explains 93.8% of the variance in mean log collapse time on the fitting data. Guided by this analysis, we propose a simple fine-tuning objective that increases . This increases Llama-3.1-8B's geometric-mean collapse time to its original value with zero additional data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.