Coarse-to-Fine Neuron-Path Discovery for Interpretable Vision Transformers
Abstract
Understanding how Vision Transformers (ViTs) compute their predictions remains challenging: most explanations highlight salient patches but do not reveal cross-layer internal information flow. Neuron-path discovery addresses this by tracing a prediction to a sequence of influential FFN neurons across layers, yet existing methods are hard to interpret (index-only neuron IDs) and expensive due to per-layer search over thousands of neurons. We present Concept-Guided Coarse-to-Fine Neuron-Path Discovery, a framework that injects human-interpretable visual concepts into neuron-path search while preserving neuron-level grounding. We first construct a global concept bank from attention-selected patch embeddings and partition FFN neurons into concept-aligned groups shared across layers. We then perform coarse-to-fine path discovery: (i) search for an influential concept-level path in a low-dimensional group-gated attribution space, and (ii) refine each selected concept group to a representative neuron to obtain a final neuron path. This design reduces the search space without sacrificing causal relevance. Experiments on ImageNet with multiple ViT variants show that our method finds neuron paths that are more causally faithful under interventions than prior neuron-path baselines, while achieving 5× practical speedups. Moreover, the resulting concept paths expose coherent semantic trajectories across depth, enabling inspection of how visual evidence emerges and where depth-localized bottlenecks arise.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.