Attribution Evaluation Needs an Ablation, Not an Annotation
Abstract
Feature attributions are evaluated by comparing what they highlight against an annotation. We show that such comparisons do not measure what they are read as measuring. On sequences where ablation demonstrates that a model uses neither the annotated feature nor the region the attribution itself identifies, attributions still locate the annotation at 5–14× the empirically measured chance level. The effect holds across 15 transcription factors, two architectures and three attribution methods, disappears when the annotated interval is compositionally rewritten, and a variance decomposition puts 40% of the score's variance on sequence identity against under 2% on the choice of method. We then measure the contaminating term rather than infer it. Scoring a model against the annotation of a factor it was never trained on sets model usage to zero by construction, and what survives is measurable: 0.22× chance without training, 1.69× after training on the benchmark, 5.45× for a byte-pair architecture. Carrying the decomposition through the argmax the metric performs identifies both terms from measured recovery alone, and the usage term is stable at 3.22±0.26 across factors whose reported scores range from 0.31 to 0.58. The architecture that appears to produce better attributions confers more recovery on annotations it never learned, and shows less usage. The protocol reproduces the effect on protein binding sites, text rationales and natural textures at 1.5–2.2×, 2.25× and 11.2×, an ordering the decomposition predicts — the protein range spanning two negative-set designs, both far below the genomic effect; images are not exempt where a composition-preserving manipulation exists. Ten corrections to the score fail, and the identification says why. A rank-based alternative improves run-to-run reproducibility on every factor in all three sequence domains and costs less than the attribution it evaluates. We derive the runs a comparison needs — four paired, fourteen unpaired — and the second kind is what this literature reports from single runs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.