Rank Is the Bottleneck: Measuring and Training Pluggable Memory for LLMs
Abstract
Pluggable memory offers a general and portable way to provide language models with reusable context: source text is encoded once into a memory matrix that can be inserted directly into the decoder's input. Yet memory is commonly characterized by its embedding count or encoding rate, which describes the allocated budget without revealing whether those embeddings provide distinct, task-useful information. We identify rank as a key constraint on this interface. For an idealized linear attention layer, the rank of the memory matrix bounds both the feature directions that attention can read and the space of score patterns across memory positions. This perspective leads to our proposed MemTower, a two-stage framework designed to make available memory directions predictive through variable-rate continuation pretraining, explicit budget announcement, rank regularization, and fixed-rate task specialization. Across multi-passage QA and reasoning benchmarks, MemTower improves over embedding-based memory baselines; at a x4 encoding rate, its 8B configuration outperforms the best baseline on all four reasoning tasks, averaging 0.79 versus 0.72. Ablations and rank interventions further show that budget announcement and regularization improve both accuracy and the use of retained directions. MemTower preserves at least 98% of its reasoning accuracy with 16 directions, whereas removing either component requires 32 despite achieving lower full-memory accuracy. Together, these results expose a distinction overlooked by memory size alone: how much representational capacity is available, and how effectively that capacity supports prediction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.