acceptodds
Under review as a conference paper at ICLR 2027

Same Answer, Different Futures: Judging Latent States by Their Continuations in Looped Transformers

Abstract

Neural network states are often judged by their current answer: a state moved into another input earns credit if the output stays correct, and computation stops once outputs settle. For a latent state updated step by step, a correct answer need not last. We study looped transformers trained on finite-group composition, a state-tracking task with provably answer-preserving edits, so states can be transplanted across edits. In every model that passes our accuracy gate, transplants decoding the correct answer turn wrong within several updates far more often than untouched runs, then almost always recover. In models whose errors come later for edits farther from the answer, the transplant's difference reaches the readout mainly through intermediate positions, and delaying it midway delays most errors about as much. This timing gives an edit-aware recipe for reusing latent states: reuse the unedited prefix exactly, warm-start the rest, and count stable outputs toward an exit only after the edit's difference reaches the readout. In those models, the recipe uses less computation than recomputing and, unlike prefix caching alone, exits sooner on average, without more wrong stops. Latent states in these looped transformers should be judged by their continuations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.