Ground, Relate, and Match: Optimal-Transport-Guided Graph Grounding for Compositional Image–Text Matching
Abstract
Dual-encoder vision-language models such as CLIP achieve strong zero-shot image-text matching, but global similarity can overlook fine-grained compositional structure, including entity binding, attributes, and directed semantic relations. Existing training-free methods introduce local evidence, yet often score components independently or reason directly over fractional alignments that do not represent coherent entity groundings. We introduce LEM (Local Evidence Matters), a training-free framework for compositional image-text matching through structured graph alignment. LEM parses paired captions into canonical directed semantic graphs and constructs a caption-neutral region graph from localized image proposals. A dustbin-augmented entropic partial optimal transport problem identifies promising entity-region correspondences while allowing entities and regions to remain unmatched. LEM converts the resulting transport plan into a compact set of injective grounding hypotheses and reweights them using the original entity-identity evidence. It then evaluates typed entity properties and directed relations over each coherent grounding and marginalizes their semantic evidence over the same support for both captions. Candidate-conditioned comparisons distinguish competing attributes, predicates, and argument directions, while independent absolute scores remain comparable across images.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.