acceptodds
Under review as a conference paper at ICLR 2027

Ground, Relate, and Match: Optimal-Transport-Guided Graph Grounding for Compositional Image–Text Matching

Abstract

Dual-encoder vision-language models such as CLIP achieve strong zero-shot image-text matching, but global similarity can overlook fine-grained compositional structure, including entity binding, attributes, and directed semantic relations. Existing training-free methods introduce local evidence, yet often score components independently or reason directly over fractional alignments that do not represent coherent entity groundings. We introduce LEM (Local Evidence Matters), a training-free framework for compositional image-text matching through structured graph alignment. LEM parses paired captions into canonical directed semantic graphs and constructs a caption-neutral region graph from localized image proposals. A dustbin-augmented entropic partial optimal transport problem identifies promising entity-region correspondences while allowing entities and regions to remain unmatched. LEM converts the resulting transport plan into a compact set of injective grounding hypotheses and reweights them using the original entity-identity evidence. It then evaluates typed entity properties and directed relations over each coherent grounding and marginalizes their semantic evidence over the same support for both captions. Candidate-conditioned comparisons distinguish competing attributes, predicates, and argument directions, while independent absolute scores remain comparable across images.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.