acceptodds
Under review as a conference paper at ICLR 2027

Modality-Assignment Sensitivity in Universal Multimodal Retrieval

Abstract

In mixed-modality retrieval, fixing candidate identities and modality counts still leaves open which entities are represented as images or text. We study sensitivity to this modality assignment across 16 representative universal multimodal embedding models across three image–text datasets. For each pool, the query, instruction, positive, and negative identities are fixed. First, we increase the number of hard negatives represented in the modality opposite to the positive while decreasing that number among controls, preserving the total image and text counts. Second, we keep each negative group half image and half text and enumerate all assignments of which entities appear in each modality. Retrieval shows assignment-sensitivity under both settings. We conduct analyses and find that rank sensitivity is concentrated among a few competitors. A mean-shift approximation built on the original scores in the positive modality reproduces a substantial part of both coverage-path endpoint changes and fixed-group-count sensitivity despite candidate-level score errors. An audit further reveals attainable reversals in model comparisons that could be missed by a specific assignment. These findings identify modality assignment as an evaluation variable, alongside candidate identities and modality counts, in mixed-modality retrieval.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.