acceptodds
Under review as a conference paper at ICLR 2027

Rank Is the Bottleneck: Measuring and Training Pluggable Memory for LLMs

Abstract

Pluggable memory offers a general and portable way to provide language models with reusable context: source text is encoded once into a memory matrix that can be inserted directly into the decoder's input. Yet memory is commonly characterized by its embedding count or encoding rate, which describes the allocated budget without revealing whether those embeddings provide distinct, task-useful information. We identify rank as a key constraint on this interface. For an idealized linear attention layer, the rank of the memory matrix bounds both the feature directions that attention can read and the space of score patterns across memory positions. This perspective leads to our proposed MemTower, a two-stage framework designed to make available memory directions predictive through variable-rate continuation pretraining, explicit budget announcement, rank regularization, and fixed-rate task specialization. Across multi-passage QA and reasoning benchmarks, MemTower improves over embedding-based memory baselines; at a x4 encoding rate, its 8B configuration outperforms the best baseline on all four reasoning tasks, averaging 0.79 versus 0.72. Ablations and rank interventions further show that budget announcement and regularization improve both accuracy and the use of retained directions. MemTower preserves at least 98% of its reasoning accuracy with 16 directions, whereas removing either component requires 32 despite achieving lower full-memory accuracy. Together, these results expose a distinction overlooked by memory size alone: how much representational capacity is available, and how effectively that capacity supports prediction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.