acceptodds
Under review as a conference paper at ICLR 2027

Lost in Ranking: When Attribution Scores Hide Class Contrast

Abstract

Attribution methods assign spatial importance for a requested class, but scoring each class map separately can conceal distinctions already present in their difference. We show that this loss changes model comparisons and derive an exact decomposition of its two sources: spatial content common to both maps alters their individual rankings but cancels on subtraction; maps highlighting different objects supply complementary evidence when the evaluated regions include background. In vision-transformer experiments on structured and randomly sampled image pairs, contrast scoring raises gradient–attention rollout’s mean source AUC from 0.647 to 0.814 and from 0.713 to 0.891, respectively. A supported advantage of one model under separate-map scoring becomes unresolved under contrast scoring in both populations. Common-component removal accounts for 69–78% of the aggregate scoring gaps for rollout and Transformer Attribution. We also prove that any nonconstant contrast-only score on unrestricted maps must depend on the maps’ relative amplitude. Mass, sign, and spatial-support controls establish where the conclusions transfer. The resulting evaluation principle is concrete: retain the maps, declare their amplitude convention, and form the class contrast before reducing it to a spatial score.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.