Routing the Past: How Retrieval Circuits Control Apparent Memory Capacity
Abstract
Ask a language model for the current value of a variable that has been reassigned times and it answers at chance. This is usually read as working memory running out. We show that much of this apparent capacity is set by which query-time circuit retrieves the answer, not by what the model stores. The evidence is causal. An answer-agnostic scalar gain on a few copy heads, applied only at the query token so that it cannot alter the context, flips genuinely binding-limited base models from failing to passing (Llama-3.1-8B , Pythia-6.9B ; positive in of models at one fixed setting; budget-matched random heads reach only against ), and ablating the same heads runs the switch in reverse. The lift survives free generation and interleaved distractors that remove the “last line” cue. The value was in the residual stream all along: under adversarial controls, a linear probe reads it above the model's own output in of models across seven families. The mechanism transfers to code, dialogue, and documents (code retrieval with no code-specific labels), and it reframes evaluation: under a raw-completion harness the capacity ranking of chat models dissolves and the scale trend vanishes. Across models (70M–32B) and three benchmarks (PI-LLM, bAbI, MultiWOZ), we characterize the occurrence-order code that crowds with load, prove that position-only routes can reach only endpoints, and identify a query-conditioned pointer-head signature behind the rare interior access of gemma-4-31B. Apparent working-memory capacity is, in large part, retrieval routing.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.