Calibrating Causal-Localization Scores: Feature-space comparisons depend on intervention budget and write-back semantics
Abstract
Causal-localization benchmarks compare representations by intervening on selected features and measuring the resulting counterfactual behaviour, and the scores are often read as properties of the representations themselves. We argue that this reading is incomplete: the measured effect is jointly determined by the representation and by the intervention used to interrogate it. We therefore treat causal localization as an intervention-relative measurement. Instead of assigning each feature space one intrinsic score, we characterise an intervention-response profile: causal response as a function of the residual-space intervention actually realized, together with its selection, write-back and evaluation conditions. Across RAVEL, MIB's IOI setting and a synthetic model with known causal structure, this distinction is consequential. Method orderings change across effective residual-space budgets, and rank-matched random subspaces acquire substantial causal effect. Nominally comparable learned interventions can occupy residual subspaces whose ranks differ by up to in our replay of MIB's learned-mask procedure, and re-optimising at a common residual-support rank changes the ordering in both successfully optimised large-mismatch tests. Write-back semantics add a separate asymmetry: changing only the operator substantially changes an SAE's causal score while leaving orthonormal bases invariant, and swapping an entire dictionary recovers 7% to 99% of the corresponding full-state causal effect across the settings we measure. These results motivate a calibration principle: compare interventions in the residual space where they act, report the intervention alongside the score, and distinguish localization from the capacity and transmission of the measuring intervention.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.