Rambling on Repeat? Exact First-Exit Laws and Minimal Intervention for Verbatim LLM Loops
Abstract
Modern LLM generations occasionally enter verbatim repetition and run to their output caps, wasting user token budgets, increasing request latency, exhausting computing power, and exposing a system vulnerability. Existing repetition penalties and -gram bans are global heuristics that distort healthy text but neither estimate a detected loop's persistence nor certify exit. We combine serving-stack signals and rigorous first-exit theory to quantify and control this behavior with guarantees on minimum intervention cost. Temperature- decoding samples the Boltzmann distribution over token energies, making the runner-up gap an activation energy. Hazard products along one greedy trace give the exact finite-horizon distribution of exit time and phase at any temperature. If a measured lap repeats indefinitely, escape obeys an Arrhenius law governed by the cycle's minimum gap, with exits concentrating at that weakest phase. The parameter-free law is validated across 37 orbits from six families up to a 284B MoE model including its 3-bit quantization. Inverting it yields two training-free, event-triggered controllers for a user-specified risk tolerance : a provably minimal-temperature pulse that exits within a chosen horizon, and a stepwise minimum-KL guard that contains every detected trajectory within tokens at minimum compatible distortion, each with successful-exit probability at least . Deployed as a vLLM logits processor, the untuned guard raises Qwen3-4B sustained escape from 0.71 to 0.91 and reduces cap exceedance on 710 matched natural-workload prefixes from 30.7% to 5.1%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.