acceptodds
Under review as a conference paper at ICLR 2027

When Speed Costs Accuracy: Causal Faithfulness of Component Attribution for Gender-Bias Localization

Abstract

There are mechanistic techniques that aim to compute the effect of individual components or groups of components on a certain type of output. Moreover, these methods are widely used to target the components that produce stereotypical behavior. Sometimes we want affordable solutions instead of generic tasks, such as IOI, Greater-Than, for quick results – and for this, we need to understand whether cheap surrogates (e.g., DLA, EAP-IG) actually recover causal importance for bias specifically, and how much of this behavior is caused by the indirect effect that is undetectable through these techniques. In our work, we benchmark four attribution methods: DLA, AtP, EAP-IG, DE against activation patching ground truth for gender-bias localization across three models: GPT-2 XL, Llama-3.2-1B, Gemma-2-2B under two ablation baselines on a gender-focused StereoSet variant. We found that the frozen-norm approximation is nearly free – it accounts for only under a sixth of the error, and removing it leaves every ranking metric unchanged. The indirect effects account for the whole problem – 88-110% of the total – leaving direct attribution comparable to or worse than a constant-zero predictor, and recovering only about half of patching's ten most important components. A single backward pass recovers the rest: attribution patching finds 7-9 of that top ten at three forward-equivalent passes rather than 234-1248. The difference does not survive the intervention: ablating each method's selected components at matched capability cost removes a comparable share of the measured bias regardless of which method chose them, so ranking components by individual effect does not determine how good they are as a set. Faithful single-component attribution is therefore necessary but not sufficient for localization: choosing a set of components needs a set-level objective, which none of these methods provide. We report stereotype score and language-modeling scores alongside a continuous bias measure, since a mean logit difference can be driven to zero by shrinking every logit.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.