acceptodds
Under review as a conference paper at ICLR 2027

Memory Pressure in Recurrent Models: Behavior, Mechanisms and Representations

Abstract

While you read this abstract, try to remember how many points fine-tuning adds to cued recall. This instruction is already changing how your brain registers this abstract, and arguably is a major difference between recurrent models (RMs) with finite state and Transformers, which can attend back to any context. Consequently, an RM must commit to its finite memory what it thinks will be relevant in the future and compress information on the go. This intuition behind state and data dependent memory drove recent architectural designs like Gated DeltaNet and RWKV-7. However, we find that language-trained RMs do not actually behave this way. In a task we call cued key–value recall, a cue that names the future query helps mostly through exact matching: the models barely process its meaning, and telling them to ignore a record still helps them recall it. Fine-tuning adds 89 points to RWKV-7’s cued recall (8% to 97% at 256 biographies), yet what the models learn remains largely token matching. Trained from scratch on a synthetic version of the task, RWKV-7 behaves rationally: uncued recall is limited by memory capacity, cued recall by length generalization, and a smaller state starts relying on the cue sooner. But the same limit that hurts recall can help learning: a finite state must compress, and we find that this inference-time compression pressure can actually push a model towards learning better features and concepts. When trained only to recall given pixels of a digit, an RM learns to draw the whole digit, while a Transformer trained on the same data and loss simply copies from its inputs. Finally, we show that these results compile into a phase diagram where task load and memory load both contribute to inference-time pressure, creating better-generalizing features. Now, do you remember how many points fine-tuning added? Perhaps the cue helped, even though you do not have quadratic attention.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.