acceptodds
Under review as a conference paper at ICLR 2027

ViTLens: A Lightweight Interpretability Tool for Vision Transformers

Abstract

Do Vision Transformers (ViTs) represent human-interpretable semantic concepts, and if so, where do they emerge across modules and how do they influence downstream predictions? We introduce ViTLens, an efficient, non-invasive framework for arbitrary-location concept probing and attribution from forward activations using only per concept. We then estimate each concept’s influence on the prediction using our adapted SVARM-LOO and attribute that influence across ViT's modules with MC-Shapley. ViTLens achieves competitive concept identification on CIFAR-100, CUB-200-2011, and ImageNet-1k while fitting up to 1490 faster than concept bottleneck model variants per probing site, and achieves a Pearson correlation of up to 0.95 with exact Shapley value at substantially lower cost. We further demonstrate ViTLens's capabilities in selective concept suppression with minimal collateral damage and reveal that old semantic concepts remain recoverable despite severe catastrophic forgetting.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.