When Can Attention Heads Be Statically Defined?
Abstract
Self-attention recomputes token interactions for each input, although some heads learn similar attention patterns across inputs. Reusing fixed patterns for these heads could reduce training cost by removing query-key score computation and softmax. Through controlled pretraining comparisons, we identify Selective Attention Freezing (SAF), a recipe that selects heads with low attention-pattern variance and fixes their attention weights to fitted post-softmax means halfway through training. We represent these fixed causal patterns with absolute-position and relative-distance preferences, reducing their storage from quadratic to linear in sequence length. A fused kernel reconstructs the patterns during execution and processes ordinary and replaced heads in a single launch. At matched training-token budgets, our preferred 25% replacement setting gives 1.056× faster post-replacement optimiser updates at 124M parameters and 4K context, with a 0.77% increase in perplexity. SAF transfers to 1B parameters at 8K context, giving a 1.066× update speedup with a 0.51% perplexity increase. Replacing 50% of heads extends the speedups to 1.119× and 1.165×, with perplexity increases of 2.49% and 2.03%, respectively. The resulting checkpoints also support downstream finetuning with faster updates for long inputs and faster causal prefill.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.