Probabilistic Set Representations with Partial Optimal Transport for Image–Text Retrieval
Abstract
Image–text retrieval involves incomplete semantic correspondence: an image may contain multiple entities and contextual details, while its caption describes only a subset. Capturing this structure requires both semantically distinct representations and a matching mechanism that accommodates unmentioned content. We propose a framework that independently represents each image and caption as a weighted set of von Mises–Fisher components. Each component is parameterized by a direction, a concentration, and a mixture weight, and cross-modal compatibility is computed through analytic distributional overlap. To encourage distinct semantic roles, we use limited entity annotations during training to supervise component support and cross-modal identity through permutation-invariant assignments. We then perform asymmetric partial alignment using optimal transport with an image-side dustbin, allowing a prescribed fraction of image mass to remain unmatched while preserving the relative coverage of caption components. Entity annotations are used only for training; inference requires only the full image and caption, with no entity inputs or auxiliary grounding modules. We evaluate retrieval alongside component localization, entity correspondence, and deletion-based selectivity to examine whether the learned components support meaningful partial alignment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.