acceptodds
Under review as a conference paper at ICLR 2027

RAM: Routed Associative Memories Learn to Recall in Recurrent Language Models

Abstract

Recurrent sequence mixers such as Mamba-2, Gated DeltaNet, and Kimi Delta Attention offer a more efficient alternative to softmax attention, but lag behind on in-context recall. On synthetic MQAR, an Online Hebbian MLP achieves information-theoretically optimal storage scaling and matches attention KV-cache capacity with up to fewer state scalars. Yet pretrained recurrent language models recall substantially worse than attention. To study this discrepancy, we introduce noisy MQAR, which combines in-context recall with competing next-token supervision. Using a polynomial-kernel memory family, we identify gradient interference from the background objective and greater parameter sensitivity of low-degree recall margins as two obstacles to learning recall circuits. We address these obstacles with Routed Associative Memory (RAM), which upweights recall-relevant tokens and routes computation across multiple recurrent memories. At 1.4B parameters, RAM closes 56.02% of the average-recall gap to attention with 0.51% more parameters than Gated DeltaNet, while its LM-Evals score is within 0.30 percentage points of the Gated DeltaNet baseline. At 360M parameters, adding -gram history to the router closes the average-recall gap to attention entirely with two memories and 94.32% of it with eight.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.