acceptodds
Under review as a conference paper at ICLR 2027

Decoding Dynamics Reveal Textual-Repetition Loops in Large Reasoning Models

Abstract

Large reasoning models can spend their generation budget repeating text without completing an answer. We study which information in the model's next-token predictions distinguishes these textual-repetition loops from successful reasoning. Across 25 models and three benchmarks, we compare six signals computed from the top- candidates recorded at each step: the concentration of the current prediction, and the recurrence of candidate probabilities, ranks, or identities at earlier steps that emitted the same token. For loops that have met a textual-repetition criterion, the recurrence of candidate identities, measured by set overlap, separates them from still-running successful generations most strongly, and controls show that candidate sets add information beyond the text-recurrence signals we evaluate. Accumulated with a moving average, the overlap yields an online detector that reads no hidden states and trains no classifier; its window and threshold are chosen on validation problems. Under greedy decoding, it alerts before the criterion on 85.3% of test-split loops at a 0.73% false-positive rate on successful generations, averaged over models. On nine models, prompting a final answer at its alerts reduces loops by 91.3% and generated tokens by 55.5%, while 99.89% of previously correct answers remain correct. Candidate-set recurrence is thus an observable of established repetition and a selective trigger for intervention under greedy decoding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.