acceptodds
Under review as a conference paper at ICLR 2027

Label-Agnostic Attribution for Interpretabilit

Abstract

Attribution maps help users inspect which parts of an image influence a classifier's prediction. They usually explain a chosen class, yet an image can support several plausible labels. In such cases, users may want to inspect the prediction before deciding which class to explain. We develop an attribution method that uses a common objective across all classes: a symmetric measure of how far their predicted probabilities depart from a uniform distribution. Sampled classes guide perturbations of the input, while gradients of this shared objective determine feature importance. This separates the class used to explore an image from the criterion used to explain it. Across the convolutional classifiers and confidence groups studied, restoring highly ranked image regions recovers class scores more effectively than the compared methods. Other methods achieve stronger score suppression when regions are removed. These findings support distribution level attribution as a complement to maps for individual classes, with its benefits depending on how feature importance is evaluated.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.