On the Role of Mechanistic Interpretability for Vision-Language Prompt Learning
Abstract
While mechanistic interpretability tools such as sparse autoencoders (SAEs) have been proposed for extracting monosemantic, human-understandable features from vision-language models (VLMs) such as CLIP, existing work focuses on using these tools for post-hoc diagnostic analysis. We posit that SAE-based interpretability methods are more than passive probing tools, but can actively guide adaptation of VLMs to downstream tasks. We propose IPL (Interpretability-Guided Prompt Learning), a framework which leverages SAE decoders to extract interpretable concept directions, composes them into prompt tokens via a learnable attention selector, and injects the resulting tokens into both the vision and text encoder layers of CLIP for adaptation. A projection-preservation regularizer keeps prompt-token geometry aligned with the underlying concept directions. Across 15 benchmark datasets and experimental settings covering base-to-novel generalization, domain generalization, cross-dataset transfer, and few-shot learning, IPL consistently outperforms prior prompt-learning methods, with unified (multimodal) concept directions achieving the strongest results, compared to vision-only and text-only concept directions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.