acceptodds
Under review as a conference paper at ICLR 2027

Linear Attention lacks Incentive for Associative Recall, not Capacity

Abstract

Linear-attention models exhibit disproportionate deficits in prefix-matching associative recall, particularly at long retrieval distances, even when their overall language-modeling loss matches or improves on softmax attention. Motivated by the performance gap, we revisit both real and synthetic recall tasks with different linear attention architectures, narrowly defining the recall-gap at an academic LLM scale. Consistent with prior work, we observe that this architecture-dependent gap is concentrated on precise predictions requiring prefix-matching associative recall. To investigate this gap, we use MQAR and find a shared empirical capacity frontier across the tested linear-attention architectures. Extrapolating this scaling, we find that a model matching a single layer of a trained LLM, can learn to retrieve thousands of synthetic associations across long distances. These results suggest that capacity alone does not explain the natural-language recall gap at this scale. We instead hypothesize that standard pretraining provides insufficient supervision for robust associative retrieval. A simple in-context repetition scheme improves recall-intensive benchmark performance and narrows the long-distance prefix-matching gap without changing the architecture, token budget, or number of optimization steps, at a slight cost to general language modeling. Targeted changes to the training distribution can thus elicit substantial but underutilized retrieval capacity in linear-attention models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.