acceptodds
Under review as a conference paper at ICLR 2027

SAHR: Semantics-Aware Head Reconfiguration for Efficient Task Adaptation

Abstract

Vision-language models (VLMs) such as CLIP have shown strong transferability across downstream tasks. However, efficiently adapting CLIP under limited supervision remains challenging because existing parameter-efficient fine-tuning (PEFT) methods typically modify prompts, lightweight modules, or support features without explicitly identifying which internal visual Transformer units are most relevant to a target task. In this work, we observe that attention heads in the CLIP visual Transformer exhibit semantic specialization, with different heads consistently capturing different visual or semantic attributes. Accordingly, we propose Semantics-Aware Head Reconfiguration (SAHR), an efficient framework that directly updates task-relevant attention heads during CLIP adaptation. SAHR constructs text-aligned semantic profiles for attention heads and quantifies their semantic correspondence to downstream tasks through task-aware descriptions. Based on this head-task correspondence, SAHR adapts only a sparse set of top-ranked heads. Extensive experiments on 11 few-shot classification benchmarks, Flickr30K and MSCOCO image-text retrieval, and OOD robustness show that SAHR outperforms recent PEFT methods with low parameter scale. In the 16-shot setting, SAHR achieves 83.1% average accuracy on 11 benchmarks with only 1.18M parameters, outperforming the strongest PEFT method while reducing parameters by 94.6%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.