Beyond Noisy Proxy: Mitigating the Gap for CLIP Based Source-Free Domain Adaptation
Abstract
Source-free domain adaptation (SFDA) aims to adapt a pretrained source model to an unlabeled target domain without access to source data, aligning with the increasing demand for data security and privacy. Recent advances in vision-language models (VLMs), such as CLIP, provide a promising alternative by exploiting cross-modal knowledge as proxy supervision to facilitate SFDA. However, even with pretrained CLIP to guide SFDA, performance remains sub-optimal. Beyond attributing this to noisy or task-agnostic proxy supervision, we explore a more comprehensive perspective: the modality gap between the image and text embeddings. On the one hand, this gap hinders the CLIP proxy from early and fully exhibiting task-specific distinctions for SFDA. On the other hand, under domain shift, it inevitably induces noisy pseudo-labels from the source model and proxy model. To mitigate these challenges, we propose a unified framework, termed Semantic Cross-modal reAlignment with seLective Pseudo-labeling (SCALP), which facilitates SFDA via improved CLIP proxy supervision and reliable pseudo-label learning. Specifically, we first propose a Class-centroids Distributional Modality Realignment (CDMR) to match textual embeddings with confident visual class-centroids via a training-free affine transformation, so as to mitigate the modality gap. We subsequently induce an Agreement-based Selective Pseudo-label Learning (ASPL) strategy that leverages dual-model agreement to reduce noisy pseudo-labels during adaptation. Extensive experiments demonstrate that SCALP consistently improves adaptation performance across multiple benchmarks, achieving overall state-of-the-art results for SFDA.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.