acceptodds
Under review as a conference paper at ICLR 2027

Semantics as Distributions: Structural Measurement for Vision-Language Representation Alignment

Abstract

Semantics in vision-language representation spaces is distributional: a semantic concept is expressed through a distribution of cross-modal correspondence patterns rather than a single pairwise response. How such distributional structure should enter alignment measurement, however, remains largely unexplored. Empirically, correspondence vectors collected from image-text pairs sharing the same semantic concept form reproducible and concept-dependent distributions, whose organization we term semantic structural fingerprints. Motivated by this observation, we introduce a decision-theoretic formulation that recasts cross-modal similarity measurement as discrimination between latent matched and mismatched correspondence distributions. By locally approximating the Bayes-optimal log-likelihood ratio to second order, we derive a principled measurement of correspondence strength and structural organization. Based on this formulation, we develop Distributional Structural Measurement (DSM), a computational realization that learns match- and mismatch-associated structures through signed low-rank components at complexity. Experiments across architectures and retrieval benchmarks consistently demonstrate the effectiveness of DSM, while statistical analyses validate the reproducibility and distinctiveness of semantic structural fingerprints. Together, these results support a distributional view of cross-modal semantic measurement. Code will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.