acceptodds
Under review as a conference paper at ICLR 2027

Characterizing Weight-Level Concept Representations in Vision models

Abstract

Concept-based explanations reveal which human-understandable concepts a network encodes, but mostly locate them in activation space, as directions or probes, rather than in the parameters that compute them. Mechanistic interpretability computation in weights and circuits, but typically for behaviors the model was trained to produce, such as a task, subtask or output class. We bring the two together: starting from human-defined concepts the model was never trained to predict, we learn sparse binary masks over the weights of frozen vision models that isolate the parameters needed to detect each concept. We first show that these masks are valid concept detectors, and then test whether this parameter-level structure is unique. Our main finding is that it is not: masks trained with different random seeds are equally valid but select different weights. The weights they all share often fail to detect the concept on their own, and ablating each mask changes the network's class predictions differently. A single mask is thus enough to detect a concept, but not to claim which weights compute it; such claims require several independently trained masks and a direct test of what they share.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.