acceptodds
Under review as a conference paper at ICLR 2027

Rambling on Repeat? Exact First-Exit Laws and Minimal Intervention for Verbatim LLM Loops

Abstract

Modern LLM generations occasionally enter verbatim repetition and run to their output caps, wasting user token budgets, increasing request latency, exhausting computing power, and exposing a system vulnerability. Existing repetition penalties and -gram bans are global heuristics that distort healthy text but neither estimate a detected loop's persistence nor certify exit. We combine serving-stack signals and rigorous first-exit theory to quantify and control this behavior with guarantees on minimum intervention cost. Temperature- decoding samples the Boltzmann distribution over token energies, making the runner-up gap an activation energy. Hazard products along one greedy trace give the exact finite-horizon distribution of exit time and phase at any temperature. If a measured lap repeats indefinitely, escape obeys an Arrhenius law governed by the cycle's minimum gap, with exits concentrating at that weakest phase. The parameter-free law is validated across 37 orbits from six families up to a 284B MoE model including its 3-bit quantization. Inverting it yields two training-free, event-triggered controllers for a user-specified risk tolerance : a provably minimal-temperature pulse that exits within a chosen horizon, and a stepwise minimum-KL guard that contains every detected trajectory within tokens at minimum compatible distortion, each with successful-exit probability at least . Deployed as a vLLM logits processor, the untuned guard raises Qwen3-4B sustained escape from 0.71 to 0.91 and reduces cap exceedance on 710 matched natural-workload prefixes from 30.7% to 5.1%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.