acceptodds
Under review as a conference paper at ICLR 2027

Supervised Concept Discovery for Interpreting Vision Representations

Abstract

Vision foundation models encode rich visual information, yet their representations do not directly expose human-interpretable factors. Concept discovery seeks to reveal these factors by decomposing pretrained representations into latent directions. However, reconstruction and sparsity objectives provide limited explicit guidance toward semantically coherent factors. We investigate class labels as a source of such guidance and introduce label impurity, an objective for guiding factor learning. Minimizing label impurity encourages each factor to exhibit consistent activation states among images of the same class, while allowing factors to be shared across classes. We derive an unbiased minibatch estimator for this objective to support stable stochastic optimization. Empirical results show that our approach discovers interpretable and visually coherent factors across diverse visual recognition datasets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.