Expressivity and Limitations of Projection Tying in Self-Attention
Abstract
Self-attention typically uses separate query, key, and value projections, yet the effects of tying these projections on expressivity remain poorly understood. We provide a systematic theoretical comparison of query-key (Q=K) and key-value (K=V) tying. We first prove that, for fixed sequence length and hidden dimension, Transformers with positional encodings remain universal approximators even under full tying (Q=K=V), given sufficient depth. At the level of individual attention heads, however, the two tying schemes impose distinct limitations. Under Q=K, the Gram structure of unmasked attention logits precludes arbitrarily accurate approximation of cyclic-shift attention patterns, regardless of head dimension. With causal masking, Q=K can approximate any causally admissible attention pattern, but previous-token routing requires asymptotically larger projected-feature norms than in the untied case. In contrast, K=V preserves attention-pattern realizability while coupling addressing and content transmission through a shared representation. We prove that the minimum head dimension required for K=V to exactly reproduce a given untied head equals the dimension of the joint span of that head’s effective key and value subspaces, and can possibly be twice as large as in the untied case. Controlled synthetic experiments corroborate these theoretical separations. Together, these results distinguish model-level universality from head-level expressivity and limitations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.