Towards compressive and scalable recurrent memory
Abstract
Recurrent memory unlocks an unbounded context horizon for language models within a finite memory footprint. We introduce Elastic Memory, a memory architecture that compresses historical context into polynomial states and reconstructs these features for attention. By leveraging HiPPO for parallelized online compression and polynomial sampling for flexible retrieval, our approach operates without additional trainable parameters. This design decouples memory capacity from the backbone model, maintaining a fixed-size recurrent state regardless of sequence length and model size. In from-scratch pretraining, Elastic Memory outperforms strong baselines on PPL and LongPPL across three long-document domains. Notably, it achieves lower perplexity than Memorizing Transformer with a 16 smaller memory footprint, and scales robustly with model size and memory size. During continued pretraining (CPT) on Llama-2 7B, it matches FreqKV's quality while delivering a 1.83 throughput speedup. For supervised fine-tuning (SFT), Elastic Memory achieves strong results on LongBench, including a 5.31-point gain over FreqKV on RepoBench-P. Elastic Memory provide a principled, scalable and parameter-free recurrent memory for LLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.