acceptodds
Under review as a conference paper at ICLR 2027

Probabilistic Set Representations with Partial Optimal Transport for Image–Text Retrieval

Abstract

Image–text retrieval involves incomplete semantic correspondence: an image may contain multiple entities and contextual details, while its caption describes only a subset. Capturing this structure requires both semantically distinct representations and a matching mechanism that accommodates unmentioned content. We propose a framework that independently represents each image and caption as a weighted set of von Mises–Fisher components. Each component is parameterized by a direction, a concentration, and a mixture weight, and cross-modal compatibility is computed through analytic distributional overlap. To encourage distinct semantic roles, we use limited entity annotations during training to supervise component support and cross-modal identity through permutation-invariant assignments. We then perform asymmetric partial alignment using optimal transport with an image-side dustbin, allowing a prescribed fraction of image mass to remain unmatched while preserving the relative coverage of caption components. Entity annotations are used only for training; inference requires only the full image and caption, with no entity inputs or auxiliary grounding modules. We evaluate retrieval alongside component localization, entity correspondence, and deletion-based selectivity to examine whether the learned components support meaningful partial alignment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.