acceptodds
Under review as a conference paper at ICLR 2027

SPACE: Shapley Projection Attribution for Concept Explanations

Abstract

Concept-based explanations connect a model's prediction to ideas people can name. Scored one at a time, they overcount: two concepts that point half the same way are each credited with the sensitivity they share, and the scores sum to no meaningful total. **SPACE** measures how much of the model's local gradient each group of concepts can express, and from that one game answers three questions. (1) How much of the model does the vocabulary reach? A coverage score, in closed form. (2) How does the captured sensitivity divide among the concepts? A Shapley allocation that is non-negative, sums exactly to the captured total, and credits shared sensitivity once. (3) Which few concepts should a reader keep? A subset selection on the same game, with a greedy guarantee, which the allocation does not solve. The game costs one gradient, one factorisation, and small per-coalition algebra, with no masked inputs, reference examples, or per-subset fits, at a per-coalition cost independent of the representation dimension. On ImageNet, a ResNet50 explained through activation coordinates with fifteen texture and colour concepts, the shares add up exactly and the coverage is stated first. Keeping the three concepts ranked highest, the allocation's three explain % of what the vocabulary can, against % for one-at-a-time scores. The gap holds across five refits of every concept and is confirmed by moving the network itself. Selection on the same game comes within one point of the best possible three. On a rendered benchmark with planted factors, the allocation recovers the planted factor as often as any rule. When a factor has several near-synonymous descriptions, selection finds both planted factors on every image, while every per-concept ranking, the allocation included, fills the shortlist with synonyms. SPACE turns any set of concept directions, however obtained, into one accounted explanation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.