acceptodds
Under review as a conference paper at ICLR 2027

Alignment with Different Pairs: Cross-Modal Similarity and Retrieval

Abstract

We distinguish three questions about cross-modal alignment: what structure each modality has, whether the two modalities organize designated partners alike, and whether the deployed scorer ranks those partners highly. Sensitivity to random shuffling does not establish that relational agreement orders pairings by original-pair fidelity, and correlation with recall does not establish its value for checkpoint selection. We introduce hard negative re-pairing: an external sentence encoder guides a one-to-one reassignment to semantically similar captions from other images, while representations and, in retrieval experiments, all image–text scores remain fixed. On COCO, hard negative re-pairing with no original partners retained yields higher centered kernel alignment (CKA) than the mean under random re-pairing retaining most original partners. Across 22 global image–text checkpoints, neighborhood agreement retains a larger fraction of its original-minus-random contrast than assigned-pair recall. On Flickr30k, selecting checkpoints by validation agreement rather than validation recall costs 5.75 percentage points of held-out bidirectional Recall@1, despite a positive agreement–recall correlation. Relational agreement alone does not establish original-pair fidelity, and checkpoint-selection criteria should be judged by the held-out retrieval performance of their choices.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.