Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory
Abstract
Decoder-only language models (LMs) entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independently. Memory Decoder introduces an external parametric long-term memory and shows primitive results, albeit on small scales. In this work, we present Memory Decoder at Scale and systematically study memory scaling with model and data size, parameter allocation between memory and the LM, and pretrained memory's ability to capture general and domain-specific knowledge. We scale external memory to 6.9B parameters and pretrain it on 300B tokens. At this scale, we address retrieval scalability and storage overhead through a distributed pipeline for FAISS indexing and retrieval and sparse kNN distribution storage, respectively. Our results show that pairing larger memories with smaller LMs outperforms scaling the LM alone. This architectural disentanglement of long-term memory from the LM substantially improves parameter efficiency. A 6.9B general memory enables Pythia-410M to surpass Pythia-12B on average across 17 benchmarks with 39% fewer total parameters. Domain-specific memory outperforms the strongest reported baseline among CPT, LoRA, and RAG by 4.05–8.53 points on average. Overall, our results establish the architectural advantage of external memory over conventional decoder-only LMs, paving the way toward disentangling long-term memory from the LM.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.