acceptodds
Under review as a conference paper at ICLR 2027

Tracing Repetitive Degeneration to the Post-Training Pipeline

Abstract

Large Language Models (LLMs) have achieved strong performance across diverse tasks. However, LLMs occasionally suffer from output repetition, which degrades output quality and wastes computational resources. This issue is widely observed, but difficult to reliably reproduce, leaving its underlying mechanisms underexplored. In this work, we present a controlled empirical study of how post-training choices are associated with this phenomenon. Firstly, to enable controlled analysis, we present the Self-Reinforced Loop Approach (SRLA), which amplifies output repetition and makes this otherwise sporadic phenomenon reproducible. Using SRLA, we find that models trained with recent reward-based post-training (e.g., GRPO) are more susceptible to repetition than conventional ones. This observation motivates our intuition that output repetition is closely associated with post-training choices, which we validate through cross-model comparisons. A weaker but consistent increase remains observable in models trained on GRPO-distilled data. Analyzing token-level distributions, we identify three crucial factors, including selected-token lock-in, insufficient candidate token diversity, and weak EOS activation. Motivated by these findings, we propose two mitigations that significantly reduce failure rates by an average of 86%. These findings bridge post-training and repetitive generation behaviors, and provide practical guidance for developing more robust models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.