acceptodds
Under review as a conference paper at ICLR 2027

From Recurrent Computation to Attention Memory in Hybrid Language Models

Abstract

Hybrid language models store context in recurrent state and an attention key-value (KV) cache. To understand how they use context, we need to trace how computation shapes the memory used for a later answer. Swapping the final recurrent state measures how much it affects the answer while holding the KV cache fixed. This leaves the effects of earlier recurrent computation on the KV cache unchanged. To trace these effects, we replace recurrent-block contributions during prefill, including their feed-forward sublayers in Qwen. The replacements come from a source context that specifies a different answer. We let later layers recompute both memories, then restore either one before the query. Most of the induced change in answer preference remains after restoring the original recurrent state. Across five hybrid models, retaining the modified KV cache preserves 67-93% of the mean shift in the answer-logit margin on exact retrieval. Restoring the original KV cache instead removes or weakens this effect. Layerwise tests locate the effect in the KV caches of deep attention layers. The effect also persists through complete-answer generation. Of 200 paired SQuAD passages with 2-5-token answers, both Qwen3.5-9B and Qwen3.6-27B answer both clean contexts correctly on 162 pairs. After restoring clean recurrent state, the modified KV cache preserves the complete source answer on 98.8% and 98.1% of this shared subset, respectively. These results support a compute-commit-retrieve pattern: Recurrent computation can change an answer through attention memory, even when swapping the final recurrent state has little effect on that answer.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.