SAHR: Semantics-Aware Head Reconfiguration for Efficient Task Adaptation
Abstract
Vision-language models (VLMs) such as CLIP have shown strong transferability across downstream tasks. However, efficiently adapting CLIP under limited supervision remains challenging because existing parameter-efficient fine-tuning (PEFT) methods typically modify prompts, lightweight modules, or support features without explicitly identifying which internal visual Transformer units are most relevant to a target task. In this work, we observe that attention heads in the CLIP visual Transformer exhibit semantic specialization, with different heads consistently capturing different visual or semantic attributes. Accordingly, we propose Semantics-Aware Head Reconfiguration (SAHR), an efficient framework that directly updates task-relevant attention heads during CLIP adaptation. SAHR constructs text-aligned semantic profiles for attention heads and quantifies their semantic correspondence to downstream tasks through task-aware descriptions. Based on this head-task correspondence, SAHR adapts only a sparse set of top-ranked heads. Extensive experiments on 11 few-shot classification benchmarks, Flickr30K and MSCOCO image-text retrieval, and OOD robustness show that SAHR outperforms recent PEFT methods with low parameter scale. In the 16-shot setting, SAHR achieves 83.1% average accuracy on 11 benchmarks with only 1.18M parameters, outperforming the strongest PEFT method while reducing parameters by 94.6%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.