Retrieval Gains, Similarity Losses: Controllable Specialisation for Entity-Specific Few-Shot Adaptation
Abstract
Large pretrained image encoders capture general visual similarity well, but this can be too coarse when retrieval must identify an operationally defined entity such as a specific character, logo, or product variant. We study target-specific few-shot retrieval adaptation: independently learning a lightweight head for each target from ten target and ten non-target images while keeping the encoder frozen. The adaptation improves target retrieval, but systematically degrades similarity judgement on general image pairs that were not used for adaptation. Across CLIP ViT-L/14, DINOv3 ViT-B/16, and SigLIP base-patch16-224, with both a linear head and an MLP with one hidden layer, Recall improves and semantic AUROC declines in all six conditions; visual AUROC declines in five, with CLIP–MLP as the exception. On 100 entities, the linear head raises Recall@5 from 80.00% to 86.75% on CLIP, from 55.92% to 62.25% on DINOv3, and from 82.17% to 87.33% on SigLIP. The trade-off persists across head structures, while its magnitude varies: the MLP loses about twice as much semantic AUROC as the linear head on all three encoders, whereas retrieval ordering depends on the encoder and cutoff. Interpolating frozen and adapted scores with one per-target scalar lets a user select an operating point on this specialisation–generality frontier. It does not remove the trade-off: CLIP admits useful intermediate settings, whereas DINOv3 has no non-zero setting that preserves both AUROCs at their frozen levels. Thus, the central design question is not whether to adapt, but how much generality to exchange for target-specific retrieval.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.