TransLP++: Transductive Multimodal Linear Probing for Few-Shot CLIP
Abstract
Vision-Language Models (VLMs) such as CLIP provide strong transferable representations for few-shot learning. Existing methods such as Linear Probe++ (LP++) efficiently adapt pretrained VLMs through multimodal linear probing, but remain inductive and do not exploit unlabeled query samples. We propose TransLP++, a transductive extension of multimodal linear probing that incorporates query-set information into LP++ while keeping the pretrained VLM frozen. Specifically, TransLP++ jointly optimizes the multimodal classifier and auxiliary query assignments by combining labeled support supervision, query-set mutual information, and CLIP-guided semantic regularization. We develop an alternating optimization procedure with block-wise Lipschitz analysis and data-driven initialization for efficient adaptation. The proposed method requires no modification to the pretrained VLM and retains the lightweight, black-box nature of classifier-based adaptation. Experiments on 11 few-shot classification benchmarks and multiple CLIP backbones demonstrate consistent improvements over strong inductive and transductive baselines, including LP++, TransCLIP, and TIM++. Our code is available at https://anonymous.4open.science/r/ICLR-8924-TransLP.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.