acceptodds
Under review as a conference paper at ICLR 2027

Toaster: Nonlinear Memory from Cascaded Linear Memories and Target Propagation

Abstract

Recurrent memories let language models process long sequences with a fixed-size state. Their design must balance the richness of the stored representation with the cost of updating it. Linear memories such as DeltaNet and KDA store one matrix and take one delta-rule step per token. A chunk of such steps has an exact WY form, so they train fast. On the other hand, deep memories such as TTT and Titans store a small MLP, which can hold maps that no matrix can. However, their update has no WY form. Toaster bridges the two with one fact: a gradient step on a linear layer is a delta-rule write. It builds a two-layer memory from two linear memories with an activation between them. Target propagation gives the hidden memory its values by projecting the output error back through the output memory. Given a chunk's cross-layer signals, each layer can then run on an unmodified delta-rule kernel. Our default runs this kernel on the output layer and updates the hidden layer once per chunk. Running it on both layers gives the best model in our ablation. At matched backbones and training-token budgets, Toaster has lower loss than KDA at all five tested sizes (0.2B–1.6B non-embedding parameters) on both C4 and PG19. Its C4 perplexity is also slightly lower than the attention baseline at every size. At the largest backbone, its mean accuracy on 13 three-shot tasks is , compared with for KDA. In length extrapolation, its loss stays below KDA's up to 20K tokens ( the training length). On the 145M backbone, at a hidden width matched to Titans' memory, Toaster takes at most approximately KDA's training-step time and is – faster than the tested Titans configurations at 32K context.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.