The Loop Records Where It Slipped: Verifying Execution Slips in Looped Language Models
Abstract
Looped language models expose a hidden state after each iteration at every token. We ask whether these states identify execution slips: wrong digits in otherwise largely correct arithmetic traces. On synthetic multi-digit addition with Ouro-2.6B, a linear probe localizes a wrong digit in 97.2% of eligible erroneous answers, and error evidence becomes more readable across iterations of the same weights. Used as a verifier with a simple arithmetic length check, it selects the correct answer in 95.6% of eligible completed problems from four sampled candidates, versus 89.7% for majority voting under the same filtering. On the 253-problem retry set, selective resampling reaches 92.5% at about 1.2x the nominal generation budget, from 87.0% greedy accuracy. State transplantation gives partial, donor-dependent digit correction; removing the probe direction yields none. Ouro-1.4B and Huginn provide weaker replications; negative results on GSM-hard and knowledge tasks show that useful verification requires both discriminative scores and correct alternatives. Disagreement between recomputations can train the detector without gold labels, although selection uses gold-supervised probes. In a four-step addition chain, a fixed probe gate improves end-to-end accuracy from 44.2% to 58.3%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.