DEFTS: Data-Efficient, Training-Free Few-Shot Learning for CLIP in the Text Subspace
Abstract
Contrastive Language-Image Pre-training (CLIP) achieves impressive zero-shot image classification performance by comparing image embeddings against class-specific text embeddings. While recent few-shot methods leverage labeled training samples to improve upon this zero-shot baseline, they often predominantly rely on image embeddings and underutilize the semantic structure of the text embeddings. In this paper, we introduce DEFTS, a data-efficient, training-free few-shot classification method strictly in the text subspace. We first demonstrate empirically that while raw text embeddings can be suboptimal as direct classifiers, the linear subspace they span remains highly discriminative. DEFTS therefore proposes to reduce the image-text modality gap by projecting estimated image prototypes directly into this subspace. Importantly, by expressing these prototypes as linear combinations of text embeddings, DEFTS reduces a high-dimensional estimation problem to the derivation of just a few coefficients. These coefficients are then computed via a Bayesian posterior mean, using a prior naturally centered on the original text embedding coordinates. Extensive experiments across 12 image classification benchmarks demonstrate that DEFTS significantly improves upon existing training-free methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.