acceptodds
Under review as a conference paper at ICLR 2027

Routing the Past: How Retrieval Circuits Control Apparent Memory Capacity

Abstract

Ask a language model for the current value of a variable that has been reassigned times and it answers at chance. This is usually read as working memory running out. We show that much of this apparent capacity is set by which query-time circuit retrieves the answer, not by what the model stores. The evidence is causal. An answer-agnostic scalar gain on a few copy heads, applied only at the query token so that it cannot alter the context, flips genuinely binding-limited base models from failing to passing (Llama-3.1-8B , Pythia-6.9B ; positive in of models at one fixed setting; budget-matched random heads reach only against ), and ablating the same heads runs the switch in reverse. The lift survives free generation and interleaved distractors that remove the “last line” cue. The value was in the residual stream all along: under adversarial controls, a linear probe reads it above the model's own output in of models across seven families. The mechanism transfers to code, dialogue, and documents (code retrieval with no code-specific labels), and it reframes evaluation: under a raw-completion harness the capacity ranking of chat models dissolves and the scale trend vanishes. Across models (70M–32B) and three benchmarks (PI-LLM, bAbI, MultiWOZ), we characterize the occurrence-order code that crowds with load, prove that position-only routes can reach only endpoints, and identify a query-conditioned pointer-head signature behind the rare interior access of gemma-4-31B. Apparent working-memory capacity is, in large part, retrieval routing.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.