Darwinian Competition among Transformer Attention Heads: Geometric Fitness and Functional Inheritance
Abstract
Understanding how Transformer components shape representation geometry requires causal analyses beyond changes in task performance. We introduce the Causal Competition Framework (CCF), which quantifies a unit's geometric fitness through intervention-induced changes in representation distributions, measured here by squared Wasserstein distance. We instantiate CCF as Representation-level Darwinian Head Competition (RDHC), coupling this measurement to an explicit population process with fitness-guided replication, parameter inheritance, and mutation. Recorded genealogies show branching reproduction across generations, and offspring exhibit greater functional similarity to their designated parents than to shuffled parents under mutation, supporting functional inheritance within the constructed process. Across text classification benchmarks, attention heads exhibit heterogeneous geometric influence, while progressive removal of high-fitness heads leads to nonlinear degradation in classification performance. To examine geometric compensation, we ablate one head and amplify the output of another, defining geometric rescue as a reduction in distributional distance to the intact model. We observe heterogeneous rescue effects across damaged–rescuer pairs, with some rescuer heads providing positive geometric rescue for multiple damaged heads and some damaged heads exhibiting limited rescue across candidate rescuers. Domain-wise analyses further reveal variation in rescue effects across semantic domains. These improvements concern representation distributions. Together, these results characterize heterogeneous head influence and geometric rescue, alongside evidence of replication and functional inheritance in explicitly constructed head populations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.