Balancing Memory Pathways: Analyzing and Improving Memory Utilization in LLMs
Abstract
Recurrent–attention hybrid language models (LMs), which interleave attention and recurrent layers, are increasingly used to combine the efficiency of the recurrent layers with the strong performance of attention layers. Prior work suggests that attention and recurrent layers offer complementary pathways to use past information: attention supports precise memory recall from earlier tokens, while recurrent layers support consolidation of disparate information over long spans. However, we observe that simply having access to both pathways does not mean that hybrid LMs are effectively using them. When we analyze how much the models rely on each pathway, we find a substantially greater reliance on attention than on the recurrent state. Standard supervised fine-tuning improves overall performance but does not improve how the two memory pathways are coordinated: the model becomes more reliant on information propagated by attention layers, while its use of information propagated by recurrent layers remains limited. To encourage better coordination between the two memory pathways, we add an auxiliary loss that limits attention’s access to earlier context while the recurrent state propagates through the full sequence. By backpropagating the auxiliary loss through the recurrent layers, this objective encourages the model to retain information from earlier context in recurrent state and use it alongside information accessed through attention. We observe increased use of the recurrent memory pathway and improved overall performance. Crucially, this imbalance and the benefit of our auxiliary loss generalize: they apply to multiple recurrent-attention LMs in question-answering and agentic tasks, as well as to attention-based LMs that combine different forms of memory. Together, our findings show that simply providing multiple memory pathways does not ensure their effective use, and that targeted supervision is needed to coordinate them.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.