acceptodds
Under review as a conference paper at ICLR 2027

HI-CLIP: Characterizing Functional Interactions among Attention Heads in CLIP

Abstract

As CLIP increasingly serves as the visual backbone of multimodal models, understanding its internal visual representations is important for building reliable and controllable models. Existing interpretability studies primarily focus on what concepts individual attention heads encode or what functional roles they play, while the functional relationships among heads remain underexplored, limiting the reliability of head-level interventions such as pruning and editing. Therefore, we propose HI-CLIP, a framework for interpreting inter-head functional relationships in CLIP-ViT. HI-CLIP identifies criterion-specific expert and collaborative heads, while characterizing inter-head relationships as supportive, redundant, or interfering. Experiments on ViT-B/16 and ViT-L/14 show that HI-CLIP reveals how criterion-specific visual representations are organized and formed through head interactions. Moreover, the discovered head roles and interactions improve customized retrieval without additional training and guide parameter-efficient adaptation. Overall, HI-CLIP provides a relational perspective for understanding how attention heads interact in CLIP, while demonstrating practical utility for customized retrieval and targeted fine-tuning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.