From History to State: Revisable Memory for Looped Transformers
Abstract
Looped Transformers reuse one layer stack across several passes, and each pass hands only its final hidden state to the next pass. The outputs of the blocks inside a pass are not kept, so a later pass cannot revise what an earlier block computed. Keeping these outputs in an additive memory does not allow revision either: the value retrieved for a key is a decayed sum of every value written under that key, so a block that contradicts an earlier block can only add to the earlier value. We propose Revisable Loop Memory, \method, which gives each token a fixed-size associative memory that persists across passes. At each block, the memory predicts the value for the block's key, and \method writes only the gated difference between the block's value and this prediction. The size of each write therefore depends on what the memory already holds: a block that agrees with the memory writes little, and a block that disagrees can replace the stored value for its key while leaving associations under orthogonal keys unchanged, which an additive write can do only by erasing the whole memory. The next pass reads the memory at the loop boundary. On Ouro-1.4B under matched fine-tuning, \method improves GSM8K by 2.3 points and MATH500 by 2.2 points over native fine-tuning with 0.2% more computation per token, and it outperforms additive writing with the same memory size.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.