acceptodds
Under review as a conference paper at ICLR 2027

Adapt Part of the Way: Class Text as a Prior for Few-Shot Self-Supervised Vision Models

Abstract

Few-shot classification with vision-language models is dominated by jointly trained encoder pairs such as CLIP, whose shared image-text space allows an image to be compared directly with text. Self-supervised vision encoders such as DINOv3 now give stronger visual representations than the CLIP image tower, but they have no text tower: they are adapted from the support images alone and ignore the class name, a prior about the class that does not depend on which, or how many, support images are available. We introduce TAP (Text-Anchored Partial transport), which supplies this prior to a self-supervised encoder using only few-shot supervision. TAP aligns the class-name embeddings of an independently pre-trained text encoder to the visual space with a closed-form orthogonal alignment estimated from the support prototypes, trains a lightweight residual map from support images onto their aligned class names, and applies the map only part of the way at inference: the query moves a fraction of the way toward the map's output while each class name moves the complementary fraction toward its support mean. Adapting part of the way is what makes the learned map useful. We evaluate TAP at 1 to 16 shots against linear probes and few-shot adapters developed for CLIP, on several self-supervised image encoders and several text encoders. On ImageNet, its distribution-shifted variants and the standard 11-dataset few-shot benchmark, TAP outperforms a linear probe at every shot count, by 8.4 points at 1 shot on ImageNet, with the largest gains in the low-shot regime and under distribution shift, and it does so with frozen encoders, a single short training run and no validation labels. The gain grows with the quality of the text encoder, so a frozen self-supervised vision model keeps improving as text encoders do. TAP thus gives self-supervised vision models access to the class names that their adaptation has so far ignored, at negligible inference cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.