Representation Alignment from a Few Paired Data via Distilled Optimal Transport
Abstract
Aligning two independently pretrained unimodal encoders typically demands millions of paired examples. In this paper, we study the regime where only – pairs are available, while unpaired data on either side is abundant. Our method expresses every representation by its similarities to the paired anchors, which places both modalities in a common coordinate system; matches the two point clouds with a single entropic optimal transport problem solved over the entire pool, whose marginal constraints counteract the hubness that defeats nearest-neighbor matching; and distills the resulting plan into two lightweight encoders, so that the correspondence becomes inductive. We additionally put this marginal constraint in relation with the one-step hub corrections of the bilingual lexicon induction literature. Pairing is required only for the anchors: the unpaired collections may be drawn independently from the two marginals, so that no item outside the anchor set needs to have both of its modalities observed. We evaluate on image-text and molecule-text alignment, where our method outperforms existing approaches at the same anchor budget on Flickr30k, MS-COCO and ChEBI-20. Code is available at https://anonymous.4open.science/r/AnchorDOT-B6CF/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.