Where Compositionality Hides: Shared-Content Interference in Contrastive Vision-Language Retrieval
Abstract
Contrastive vision–language models struggle with compositional retrieval, but a low retrieval score does not by itself establish that the relevant information is absent from their representations. We identify an exact, format-dependent bottleneck in the scoring geometry. Removing the mean of two normalized candidate captions from a frozen image embedding preserves every 1×2 caption decision, yet forces the 2×2 Group Score to equal the original Text Score. On CLIP ViT-B/32, this closes the full 22-point Winoground Group–Text gap, increasing Group Score from 9.0% to 31.0% and Image Score from 11.25% to 64.5%, with replication across five models and ColorSwap. An equivalent score decomposition shows that 213 of 355 raw Image-Score failures have a positive joint-matching margin but fail the cross-image offset threshold. A complementary probe-free analysis reports that the highest-magnitude 20% of coordinates carry 57.3% of absolute similarity mass but only 20.2% of the reported discriminative share on object swaps. We distinguish this basis-dependent observation from the exact projection result. Tested global corrections do not recover the paired diagnostic’s gain, while a top-K approximation recovers part of it in the evaluated setting. Together, these results establish retrieval geometry as a concrete bottleneck in paired compositional evaluation: useful discrimination is already present in frozen pooled features, although gap closure alone neither proves complete compositional representation nor establishes open-gallery retrieval improvement.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.