Pairing Graph Identifiability of Multimodal Representations
Abstract
ImageBind style models do not observe a joint tuple of all modalities. They train on pairwise collections sampled independently across pairs, for instance an image and text dataset together with an image and audio dataset. Existing identifiability theorems assume the opposite observation, in which every view of a sample is present at once. They therefore do not name the latent factors that independent pairs identify. In this paper, we prove that a latent block is consistently block identifiable from pairwise datasets if and only if the modalities that carry it induce a connected subgraph of the pairing graph. Modalities are its vertices and observed pairs are its edges. On a star, this criterion identifies the factors that both spokes share with the hub, and that intersection is the content of emergent alignment. It does not identify sharing that lives only on the unobserved spoke pair. After each edge has estimated a shared representation, residual encoder disagreement equals the effective resistance of the pairing network. Among trees, a star minimises the worst pair error when the hub carries shared signal. Once that signal vanishes, spoke to spoke error returns to the prior. The identity is tight for linear Gaussian estimators, and it remains an upper bound with a nonnegative remainder for nonlinear mixing functions. Unpaired optimal transport mixes private coordinates, whereas paired barycentres do not. A content encoder of shared factor dimension leaks no private information given that factor, while an invertible embedding of a view leaks its private coordinates. Projection heads trained by edgewise InfoNCE on frozen CLIP and Qwen3-VL features recover the ranking of star, path, and disconnected graphs, and a nonlinear mixing model recovers the spoke factor only after the missing edge is added.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.