When Speed Costs Accuracy: Causal Faithfulness of Component Attribution for Gender-Bias Localization
Abstract
There are mechanistic techniques that aim to compute the effect of individual components or groups of components on a certain type of output. Moreover, these methods are widely used to target the components that produce stereotypical behavior. Sometimes we want affordable solutions instead of generic tasks, such as IOI, Greater-Than, for quick results – and for this, we need to understand whether cheap surrogates (e.g., DLA, EAP-IG) actually recover causal importance for bias specifically, and how much of this behavior is caused by the indirect effect that is undetectable through these techniques. In our work, we benchmark four attribution methods: DLA, AtP, EAP-IG, DE against activation patching ground truth for gender-bias localization across three models: GPT-2 XL, Llama-3.2-1B, Gemma-2-2B under two ablation baselines on a gender-focused StereoSet variant. We found that the frozen-norm approximation is nearly free – it accounts for only under a sixth of the error, and removing it leaves every ranking metric unchanged. The indirect effects account for the whole problem – 88-110% of the total – leaving direct attribution comparable to or worse than a constant-zero predictor, and recovering only about half of patching's ten most important components. A single backward pass recovers the rest: attribution patching finds 7-9 of that top ten at three forward-equivalent passes rather than 234-1248. The difference does not survive the intervention: ablating each method's selected components at matched capability cost removes a comparable share of the measured bias regardless of which method chose them, so ranking components by individual effect does not determine how good they are as a set. Faithful single-component attribution is therefore necessary but not sufficient for localization: choosing a set of components needs a set-level objective, which none of these methods provide. We report stereotype score and language-modeling scores alongside a continuous bias measure, since a mean logit difference can be driven to zero by shrinking every logit.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.