The Value of Memory in Language Models: Joint Limits on Language Model Expressivity
Abstract
The value of memory in a language model depends on the model that must use it. Prompt information matters only if it is both retained and usable by the model’s fixed parameters. We study this interaction through prompt accessibility, asking which outputs a frozen model can be prompted to generate. We place a causal cut between prompt processing and generation, so all prompt-derived information available afterward must pass through the retained state. This gives a joint bound in which prompt budget, retained memory, and compatible model-side constraints all limit the same accessible outputs, with the tightest applicable constraint governing. We then sharpen the memory-side bound under additional assumptions on the update dynamics and reachable-state geometry, and show how mechanism-specific structure can yield tighter bounds, using Gated Linear Attention (GLA) as a concrete example. Controlled experiments on two frozen GLA checkpoints show that increasing retained memory moves the discovered output frontier. On GLA-1.3B, increasing the retained payload from 12 to 24 to 36 MiB raises the evaluated 50% frontier from 8192 to 10240 to 12288 tokens. Controls further show that memory usefulness depends on both prompt budget and what information is retained. A same-weight sliding-window comparison reproduces the qualitative memory-frontier response, while Mamba provides complementary evidence from fixed-size recurrent memory. Together, these results support treating prompt budget, memory, and the learned model as coupled resources, motivating memory and architecture to be matched rather than scaled independently.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.