Transforming Attention: Generating Multi-Head Attention from Learned Group Orbits
Abstract
Multi-head attention allows transformers to capture different relationships between tokens, but typically learns separate projection parameters for every head. Sharing parameters across heads offers a route to smaller attention blocks, with the central challenge of preserving each head’s ability to specialize. We introduce group-transformed multi-head attention (GT-MHA), which generates attention heads by applying learned transformations to a compact set of shared query, key, and value projections. Each head combines shared Lie-algebra generators to form its transformation through the matrix exponential or low-order Taylor approximations. Our analysis shows how heads within a shared orbit represent different interactions within common query and key subspaces. We then extend GT-MHA to multiple learned bases (a multi-orbit formulation), allowing different groups of heads to specialize in different subspaces while controlling the degree of parameter sharing. The resulting architecture preserves the standard multi-head interface and supports key–value caching at the level of shared bases. Under matched BERT pretraining and fine-tuning protocols, residual GT-MHA achieves an official GLUE score of 70.5 versus 70.3 for standard multi-head attention, using 47.2% fewer attention-block parameters. On TinyStoriesV2 character-level language modeling, GT-MHA achieves comparable validation perplexity to standard multi-head attention (1.3331 vs. 1.3355) under the same training budget, with 54.7% fewer attention-block parameters and 75% less persistent KV-cache storage. Ablations and head-diversity analyses provide evidence of learned specialization within shared bases, supporting structured parameter sharing as an effective design for parameter-efficient attention.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.