acceptodds
Under review as a conference paper at ICLR 2027

Storing Is Not Remembering: LSTM-UT and Bounded Gated Memory for Looped Transformers

Abstract

Recurrent-depth Transformers reuse one block across many steps, so information needed later must survive repeated rewriting of the hidden state. A natural remedy is to keep more history. We show that, in controlled cellular-automaton tasks, making history available is not the same as making it usable. Using RuleĀ 30, where the correct state is known at every recurrent step, we test depth extrapolation and delayed recall, the recovery of an earlier state after further computation. CoTFormer, which caches keys and values from every earlier step, extrapolates less far and recalls less accurately than a Block Universal Transformer (BUT) that keeps only its current state. Interventions show that its retained history can pull a corrected trajectory back toward failure, and that the cache block written at the requested step is neither necessary nor sufficient for recall. We introduce LSTM-UT, which adds a small, bounded, gated cell state to the shared block. Trained to depth 12, LSTM-UT keeps \(99.7%\) exact-row accuracy at depth 60, where BUT gets no row fully correct, and one checkpoint stays above \(99.95%\) at depth 1,000. It also improves delayed recall over both baselines, and the advantage largely persists at near-matched parameter counts. On these tasks, a small state under learned control proved more useful than a complete but unaddressed history. In OpenWebText2 language modelling, LSTM-UT outperforms BUT and comes within 0.6 perplexity of a parameter-matched CoTFormer, which needs up to 91% more training time per step.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.