UVA: A UNIFIED VISION-LANGUAGE ALIGNMENT FRAMEWORK FOR FINE-GRAINED ZERO-SHOT AND FEW-SHOT CLASSIFICATION
Abstract
Vision-language models such as Contrastive Language-Image Pre-training (CLIP) achieve strong zero-shot recognition, yet they remain unreliable on fine-grained visual classification, where visually similar classes differ only in small, localised parts. Recent training-free methods address this by matching selected image regions against detailed class descriptions generated by large language models (LLMs). We show that this is insufficient: on CUB-200-2011, zero-shot region matching reaches only 61.17% accuracy, and accuracy falls to 0% on the images whose regions are least well grounded in their class descriptions. We introduce Unified Vision-Language Alignment (UVA), a framework built on a frozen CLIP ViT-B/16 that combines Attention-guided Bounding-box Selection (ABS), Bi-directional Fine-grained Text Alignment Description Refinement (BiFTA-DR), lightweight visual and textual residual adapters, and Weighted Cross-Alignment (WCA). The adapters are trained with a region-to-class-prototype contrastive objective that aligns local image regions with refined class descriptions. With 16 labelled images per class, UVA improves CUB-200-2011 accuracy from 56.02% (CLIP) to 80.29 ± 0.31% while training only 1.05M parameters, approximately 0.7% of the backbone. Across Oxford-IIIT Pet, ImageNet-100, and CUB-200- 2011, the benefit of adaptation increases with the difficulty of the zero-shot task. Finally, we show that agreement between WCA and Localized Zero-Shot Learning (LaZSL), an independent optimal-transport matcher, provides an unsupervised confidence signal: predictions on which the two agree are 28–41 percentage points more accurate than those on which they disagree.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.