Text-Conditioned Orthogonal Compensation for CLIP-based Few-Shot Adaptation
Abstract
CLIP exhibits strong transferability in few-shot adaptation, yet two key challenges remain. First, existing adaptation methods are structurally diverse, while a unified geometric understanding of their underlying connections is still lacking. Second, effective adaptation should capture task-specific information while preserving the inherent category flexibility of CLIP, which is often weakened by label-space-specific adaptation. To address these challenges, we introduce a modality-gap-inspired geometric framework that interprets major CLIP adaptation strategies through the visual–classifier mismatch, and our subspace analysis further motivates complementary classifier directions beyond the principal text-classifier subspace. Based on this insight, we propose Text-Conditioned Orthogonal Compensation (TCOC), which uses a shared text-conditioned generator to produce class-wise residuals and projects them onto the orthogonal complement of the principal text-classifier subspace. This design retains the original text classifier as a semantic anchor while introducing complementary task-specific directions. Since the generator is shared across classes, TCOC can naturally extend to new label spaces. Experiments on standard few-shot classification and multiple generalization settings demonstrate its effectiveness and category flexibility. The code will be publicly released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.