acceptodds
Under review as a conference paper at ICLR 2027

How Much Attribution Mass Lands on Tokens That Cannot Matter? A Faithfulness Benchmark for Transformers

Abstract

Attribution methods for transformers are usually judged by whether their heatmaps look plausible to a human reader. We take a more direct route. We train small decoder-only transformers on four algorithmic tasks (sparse parity, indexed retrieval, associative recall, and modular addition with distractors) in which every input position is, by construction, either causally relevant or provably unable to affect the output. Any attribution mass landing on a provably irrelevant position is an error, with no threshold, reference model, or human judgment involved in that verdict. Under one pre-registered protocol we evaluate fourteen scoring rules, and every number carries a -based confidence interval over 8 to 11 independently trained seeds. Our most surprising finding concerns training dynamics: through the grokking transition on modular addition, the faithfulness of attention-based methods degrades sharply and becomes far more seed-dependent, while test accuracy rises to 100%. Explanations of the better model are worse. Second, cheap attribution and causal measurement diverge exactly where feature interactions matter: on sparse parity a resample-ablation oracle reaches 0.031 false-positive mass while every attribution method stays above 0.49, a gap that rank-preserving sharpening cannot close, so the failure is misidentification rather than diffuse scoring. Third, protocol choices can outweigh the choice of method: on the three tasks with first-order structure, moving integrated gradients from a marginal-mean baseline to a zero baseline raises false-positive mass by factors of 2 to 22. Fourth, paired tests show the top two methods are statistically indistinguishable on two of four tasks, so single-seed rankings among mid-ranked methods carry little information. We release the task generators, model zoo, attribution harness, and evaluation suite.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.