Delta-Rule Memories Store More Than They Recall, and a Load Curriculum Closes the Gap
Abstract
Delta-rule layers are now common in open frontier models, where instead of a growing KV cache (as in attention), they keep a fixed matrix per head and use it as an associative memory. This rank matrix can return at most exact associations, but a language model only needs the correct token to win the argmax, and recent theory shows that an offline memory then returns the right token for associations whenever . We ask how many associations these models recall in practice, in two-layer, single-head DeltaNet models trained on multi-query associative recall. Using state reads and edits built from the model's own primitives, we identify a circuit in which the first layer carries each key token to its paired value and the second layer stores the associations. We then distinguish failures of writing from failures of retrieval. In width-128 models trained on 64 pairs, recall falls to at a load of . Reading the same state with each association's write key recovers , and a matrix built by ridge regression from the same keys and values raises this to , so both online writing and retrieval limit recall, and many associations remain accessible after the model fails to recall them. Finally, a training curriculum that progressively increases the number of pairs (the load) raises the load at which the model reaches recall from to at , and reaches it faster than training directly at the target load. Effective recall capacity therefore depends on the learned writing and retrieval procedures and can be increased through training without enlarging the memory state.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.