acceptodds
Under review as a conference paper at ICLR 2027

A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets

Abstract

Linear-attention and state-space language models compress the prefix into a fixed-size recurrent state, which makes them efficient but leaky: once many key–value associations compete, earlier facts are overwritten and exact recall of distant tokens fails. The brain faces the same trade-off and resolves it with complementary learning systems: a slowly consolidating neocortex compresses experience, while a fast hippocampal store keeps the surprising episodes the neocortex cannot absorb. Many efficient hybrids already give the recurrent state such a store — a small softmax-attention cache that holds exact key–value pairs — but fill it with the most recent tokens, as if the state could not tell what it has lost. We show that it can: in a delta-rule model, the magnitude of the update a token commits to the state measures exactly how poorly that token fits the compressed memory, and it is already computed at every step. HOLA (Hippocampal Linear Attention) adds to every Gated DeltaNet layer a small exact key–value memory that keeps the top-scoring tokens by this signal, however far back they are. The mechanism is visible directly: on a 32k-token sequence, the trained model's 64-slot memory still holds the needle 29k tokens later, whereas a recency cache of the same size has long dropped it. At 1B parameters trained on 50B tokens, HOLA surpasses published Gated DeltaNet, KDA, and their preconditioned variants on in-context retrieval (four-task average 35.4 vs. at most 32.7) while staying on par with them in language modeling and commonsense reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.