Inherent Trade-offs when Enforcing Correct Data Symmetries in Multi-head Self-attention
Abstract
We give a complete characterization of equivariant multi-head self-attention (MHSA): if an MHSA layer is equivariant to a symmetry group , then can only act by permuting head-clusters, with QK and OV matrices satisfying an equivariance constraint tied to the group action. As a consequence, we prove that any fixed MHSA architecture that achieves exact equivariance by polynomially parameterizing unconstrained MHSA parameters inevitably leads to expressivity loss within the class of equivariant maps: the equivariance locus of unconstrained MHSA forms a union of extremely many Zariski-irreducible components in a reduced parameter space, and any single architecture covers at most one. For acting on copies of the regular representation as the token feature space, we show that there are components for eight attention heads. We illustrate this with experiments on artificial tasks and show that the predicted head-cluster structures can arise naturally when the attention layer is pressured to become equivariant.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.