acceptodds
Under review as a conference paper at ICLR 2027

On the Role of Mechanistic Interpretability for Vision-Language Prompt Learning

Abstract

While mechanistic interpretability tools such as sparse autoencoders (SAEs) have been proposed for extracting monosemantic, human-understandable features from vision-language models (VLMs) such as CLIP, existing work focuses on using these tools for post-hoc diagnostic analysis. We posit that SAE-based interpretability methods are more than passive probing tools, but can actively guide adaptation of VLMs to downstream tasks. We propose IPL (Interpretability-Guided Prompt Learning), a framework which leverages SAE decoders to extract interpretable concept directions, composes them into prompt tokens via a learnable attention selector, and injects the resulting tokens into both the vision and text encoder layers of CLIP for adaptation. A projection-preservation regularizer keeps prompt-token geometry aligned with the underlying concept directions. Across 15 benchmark datasets and experimental settings covering base-to-novel generalization, domain generalization, cross-dataset transfer, and few-shot learning, IPL consistently outperforms prior prompt-learning methods, with unified (multimodal) concept directions achieving the strongest results, compared to vision-only and text-only concept directions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.