Alignment with Different Pairs: Cross-Modal Similarity and Retrieval
Abstract
We distinguish three questions about cross-modal alignment: what structure each modality has, whether the two modalities organize designated partners alike, and whether the deployed scorer ranks those partners highly. Sensitivity to random shuffling does not establish that relational agreement orders pairings by original-pair fidelity, and correlation with recall does not establish its value for checkpoint selection. We introduce hard negative re-pairing: an external sentence encoder guides a one-to-one reassignment to semantically similar captions from other images, while representations and, in retrieval experiments, all image–text scores remain fixed. On COCO, hard negative re-pairing with no original partners retained yields higher centered kernel alignment (CKA) than the mean under random re-pairing retaining most original partners. Across 22 global image–text checkpoints, neighborhood agreement retains a larger fraction of its original-minus-random contrast than assigned-pair recall. On Flickr30k, selecting checkpoints by validation agreement rather than validation recall costs 5.75 percentage points of held-out bidirectional Recall@1, despite a positive agreement–recall correlation. Relational agreement alone does not establish original-pair fidelity, and checkpoint-selection criteria should be judged by the held-out retrieval performance of their choices.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.