Attribution Evaluation Scores Depend on the Network Being Explained
Abstract
Attribution methods are often compared using faithfulness scores based on interventions and localisation scores based on human annotation. We show that both can depend strongly on the network being explained, even when the attribution method is unchanged. In fixed-width CIFAR-10 ResNets, rank-based faithfulness tracks how concentrated the measured behaviour is across channels (within-seed Spearman ). An extreme-weight-decay model with one effective channel reaches only 18% accuracy yet receives a score of 1.000, because almost all channel–image pairs are tied at zero attribution and zero ablation effect. On ImageNet, using the same attribution and the same images, faithfulness rises from 0.595 to 0.880 across the first three stages of ResNet-18 and reaches 1.000 at the final stage, where the linear readout makes the metric exact by construction; activation magnitude alone scores 0.887 there. In a separate localisation study, two publicly available diabetic-retinopathy classifiers with the same ViT-B/16 architecture differ by in lesion localisation, more than the variation across five attribution methods within either classifier; one falls below a matched random baseline under every method. Gain-over-random normalisation, reproducibility normalisation, and support restriction do not remove the concentration effect. We therefore recommend reporting the explained layer and its distance from the readout, the concentration of the measured effect, and results across more than one layer and model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.