Causal Darwinian Head Selection: A Mechanistic Framework for Interpreting Attention Heads
Abstract
Understanding how internal components of Transformer models contribute to prediction remains a central challenge for mechanistic interpretability. We propose Causal Darwinian Head Selection (CDHS), a Darwinian, causal intervention framework for analyzing attention heads as a population of computational units. CDHS uses temporary ordered ablations to estimate head-level causal contribution and to measure the resulting changes in model behavior; it is intended as a post hoc diagnostic method. Applied to Transformer-based image classifiers, CDHS jointly characterizes three aspects of head organization: sparse causal fitness, class-conditional behavioral effects, and checkpoint-wise emergence under task-aligned supervision. We find that a small subset of dominant heads contributes disproportionately to predictive performance, whereas temporarily intervening on many low-fitness heads has limited immediate effect. Interventions on dominant heads induce structured class-level shifts in prediction behavior, suggesting that some heads support class-discriminative computations rather than only global accuracy. Using checkpoint-wise analysis and a matched label-permutation control, we further observe that fitness concentration strengthens during training under real labels more than under permuted labels. These results position CDHS as a framework for linking causal intervention, class-level model behavior, and training dynamics in vision Transformers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.