Memory Pressure in Recurrent Models: Behavior, Mechanisms and Representations
Abstract
While you read this abstract, try to remember how many points fine-tuning adds to cued recall. This instruction is already changing how your brain registers this abstract, and arguably is a major difference between recurrent models (RMs) with finite state and Transformers, which can attend back to any context. Consequently, an RM must commit to its finite memory what it thinks will be relevant in the future and compress information on the go. This intuition behind state and data dependent memory drove recent architectural designs like Gated DeltaNet and RWKV-7. However, we find that language-trained RMs do not actually behave this way. In a task we call cued key–value recall, a cue that names the future query helps mostly through exact matching: the models barely process its meaning, and telling them to ignore a record still helps them recall it. Fine-tuning adds 89 points to RWKV-7’s cued recall (8% to 97% at 256 biographies), yet what the models learn remains largely token matching. Trained from scratch on a synthetic version of the task, RWKV-7 behaves rationally: uncued recall is limited by memory capacity, cued recall by length generalization, and a smaller state starts relying on the cue sooner. But the same limit that hurts recall can help learning: a finite state must compress, and we find that this inference-time compression pressure can actually push a model towards learning better features and concepts. When trained only to recall given pixels of a digit, an RM learns to draw the whole digit, while a Transformer trained on the same data and loss simply copies from its inputs. Finally, we show that these results compile into a phase diagram where task load and memory load both contribute to inference-time pressure, creating better-generalizing features. Now, do you remember how many points fine-tuning added? Perhaps the cue helped, even though you do not have quadratic attention.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.