The Art of Remembering: Scaling Laws for Transformer Associative Memories
Abstract
Memory is a fundamental component of intelligent systems. Classical neural memories using Hopfield networks showed that retrieval is best done by energy minimization, where a partial or noisy query is iteratively refined until it settles on a stored memory. However, modern architectures largely ignore energy minimization and instead retrieve stored memories primarily through feed-forward propagation, raising the question of whether energy-based retrieval still matters once neural memories are scaled. To answer this, we compare four retrieval processes: a single forward pass in Feed-Forward Transformers, a diffusion process in Diffusion Transformers, iterative refinement in Recurrent Transformers, and energy minimization in Energy-Based Transformers (EBTs). For each architecture, we train Transformer associative memories and characterize scaling laws and trends across three fundamental properties of neural memories: capacity, generalization to corrupted queries, and continual learning. Across all three properties, we find that EBTs consistently scale best, and that their benefits grow with scale. For example, EBTs have a capacity scaling exponent over 40% higher than other architectures, and store 3.5x more memories than the next-best architecture at scale. Beyond capacity, EBTs generalize better than models 8x larger, and continually learn new memories over 5x faster than other Transformer models. Together, these results show that the classical advantages of energy-based retrieval hold for modern Transformers at scale, offering guidance for designing memory systems in language models, world models, and other architectures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.