Don't Repeat Yourself: Stopping Verbatim Loops at Sampling Time
Abstract
Large Language Models generate text by sampling tokens autoregressively, but open-ended generation is prone to verbatim looping, where the model repeats spans already present in its context. Standard defenses, such as repetition, presence, and frequency penalties, and n-gram blocking, act on token recurrence rather than on the sequential structure of a loop, and suppress looping only at strengths that also degrade formatting and fluency. We propose Don’t Repeat Yourself (DRY), a sampling-time logit adjustment that penalizes a candidate token only when generating it would extend the current suffix into an exact continuation of a span seen earlier in the context, with sequence breakers that protect chat templates and formatting tokens. Our experiments across models from 1.5B to 120B parameters, nine prompt families, and a 600-pair human study show that DRY reduces the suffix-extension rate by 47% while improving lexical diversity. An intervention-matched placebo produces no such reduction, identifying suffix-matching as the operative mechanism. On AWQ-quantized 70B and 120B models, DRY reduces the loop rate by roughly half while preserving MT-Bench, MMLU, and GSM8k, where standard alternatives lose measurable ground. DRY has been adopted by popular open-source LLM inference frameworks, including llama.cpp, ExLlamaV2, and text-generation-webui, highlighting its impact on practical text generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.