acceptodds
Under review as a conference paper at ICLR 2027

PrefixRLVR: Mitigating Catastrophic Repetition in Long-Context RAG

Abstract

Abstract Modern end-to-end agentic workflows rely on robust long-context generation for multi-step autonomous tasks. However, sustained long-form outputs make LLMs prone to catastrophic repetition, which is particularly evident in open-source models with smaller parameter sizes, resulting in wasted compute and stalled downstream pipelines. Prior work mitigates repetition through decoding strategies, attention interventions, or training-based methods without assessing their impact on general capabilities, leaving the risk of catastrophic forgetting unaddressed. In this work, we introduce De-Rp, accompanied by a character-level detector that precisely localizes the onset of catastrophic repetition. Through systematic analyses across a range of open-source models, we find that catastrophic repetition is not merely a random decoding failure but can be induced by specific trigger prefixes that are shared across models of different series and scales. Hidden-state analysis reveals a repetition-associated signal near the prefix boundary that is shared across models. Building on this insight, we propose PrefixRLVR, a reinforcement-learning-based approach that combines training from these high-risk prefixes with length-aware advantage scaling. Experimental results show that PrefixRLVR reduces catastrophic repetition rates by approximately 49–330× across the evaluated backbones while preserving general model capabilities. Further long-context evaluations demonstrate consistent generalization gains, suggesting that PrefixRLVR offers an effective path toward more reliable long-context generation. Code and data are released at: https://anonymous.4open.science/r/PrefixRLVR.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.