acceptodds
Under review as a conference paper at ICLR 2027

CoverCLIP: Coverage-Aware Visual and Textual Alignment

Abstract

Contrastive Language–Image Pre-training (CLIP) performs strongly on image–text retrieval and zero-shot classification. However, existing CLIP training objectives treat all semantically correct captions as equally valid, largely ignoring their semantic coverage. Ideally, for a given image, a caption covering more visible semantic information should receive a higher image–text similarity score. To explicitly measure this, we introduce CoverageBench, a manually verified benchmark containing images derived from Visual Genome. Each image is annotated with four visual semantic atoms (e.g., object, attribute, relation, etc.) and paired with four captions with increasing semantic coverage, denoted as \(C_1, C_2, C_3, C_4\). Models are expected to assign similarity scores satisfying \(s(I, C_1) < s(I, C_2) < s(I, C_3) < s(I, C_4)\). Evaluating on CoverageBench, we find that existing CLIP models struggle to consistently distinguish correct captions according to their semantic coverage. To address this limitation, we propose CoverCLIP, a coverage-aware training framework built on three components. First, we construct Coverage-1M, which consists of 710K samples with graded semantic coverage captions, 228K samples with long-short captions, and 76K samples with highly descriptive captions, providing explicit supervision for graded semantic coverage. Second, we introduce a step-wise mixed training strategy that incorporates semantic coverage-aware supervision into standard contrastive learning. Finally, we design coverage-ranking and hard-negative separation objectives to further enhance coverage-aware supervision. Experiments demonstrate that CoverCLIP improves CoverageBench performance while remaining competitive on image–text retrieval and zero-shot classification. Interestingly, when applied to text-to-image generation by replacing the CLIP text encoder in SDXL, CoverCLIP transfers its coverage awareness to the generation process, enabling the model to progressively generate more detailed images as semantic coverage increases.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.