acceptodds
Under review as a conference paper at ICLR 2027

Discovering Sparse Concept Graphs as Interpretable Surrogates for Vision Models

Abstract

Recent methods for mechanistic interpretability of vision models uncover interpretable concepts and trace the circuits underlying predictions for individual samples or classes. However, predictions for different classes often depend on shared concepts, and this shared structure is central to understanding how the network operates as a whole. We address this gap by learning sparse rules that describe how combinations of concepts in one layer give rise to concepts in the next, and compose these rules across layers into a ConcepTome, a single directed concept graph that serves as a global, symbolic surrogate of the network. The ConcepTome can be examined as a whole or queried for individual samples, single classes, or groups of classes to reveal the computational structure they share. We evaluate our approach on CLIP ViT and CLIP ResNet and show that the recovered ConcepTomes remain sparse while faithfully preserving model computation. Inspection reveals how concepts are combined, reused, transformed, and specialized throughout the network. Interventions show that the recovered structure isolates a compact set of concepts on which the computations and predictions of the model depend. To make this structure accessible, we develop an interactive interface for exploring the ConcepTome, visualizing concepts and rules, and intervening on predictions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.