A Probabilistic Account of Associative Memory of a Single Attention Head
Abstract
A single attention head that predicts the next token by attending to an earlier, related position in the sequence can also be seen as a form of associative memory. However, the reason and amount of such association storage are generally unclear. In the proposed work, we investigate a non-asymptotic explanation for a single-head, attention-only transformer. We introduce random spherical representations to construct an explicit query-key matrix and a value-output matrix that jointly bind a designated target position in each stored sequence to a prescribed output token. The proposed system allows arbitrary queries and gives each sequence-position pair its own representation, unlike constructions that fix the query to an end-of-sequence token or assign a single embedding per vocabulary item, thereby allowing a repeated token to carry different associations at different positions. Our analysis of the proposed system isolates three effects, namely, query-key retrieval interference, decoder interference, and the operator-norm normalisation of the value-output map, and controls them using sub-Gaussian and sub-exponential concentration, together with a matrix Bernstein bound. It has been proved that with high probability, the construction simultaneously stores associations whenever the sequence length and vocabulary size are polynomial in head dimension, a quadratic-over-logarithmic lower bound on the associative capacity attributable to attention alone. Our experiments confirm the predicted growth of the normalisation constant's scaling exponent and a sharp memorisation phase transition. Our experiments show capacity follows the ambient head dimension when the query/value maps are learned, and shifts toward a spectral dimension when the attention geometry is fixed and strongly anisotropic.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.