Balancing Cross-Modal Semantic Alignment and Instance-Level Visual Discrimination by Controlling Constraint Density in Vision-Language Contrastive Distillation
Abstract
Vision-language models trained on large-scale image–text pairs achieve strong cross-modal semantic alignment and have become the dominant approach for cross-modal retrieval. Despite substantial progress in CLIP and its recent variants, they remain limited in the instance-level visual discrimination required for image-to-image retrieval. We first show that distilling visual discrimination knowledge from instance-oriented teachers introduces a trade-off between cross-modal semantic alignment and instance-level visual discrimination. We then identify the density of transferred neighborhood relations (constraint density) as a key factor governing the strength of this trade-off, providing a unified perspective on representative feature-, similarity-, and distribution-based distillation objectives. Based on this principle, we propose Teacher-guided Nearest-neighbor Contrastive Distillation (TNCD) as a simple realization of the lowest-constraint-density regime. Extensive experiments demonstrate that lower constraint density consistently leads to a more favorable trade-off, with TNCD serving as a practical realization that naturally supports heterogeneous teachers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.