acceptodds
Under review as a conference paper at ICLR 2027

From Controlled Edits to Complex Bindings: Evaluating and Improving Structure-Aware Vision–Language Models

Abstract

CLIP is effective at global image–text matching, and Structure-CLIP further improves discrimination on template-based attribute and relation perturbations. However, whether these gains extend to descriptions containing multiple entities, modifiers, and relations remains unexplored. A structural error in such descriptions may concern not only attribute or relation direction but also host assignment or spatial layout. We construct 666 image–text instances over six structural categories. Each instance pairs a natural image with a complex multi-entity caption P1, a meaning-preserving paraphrase P2, and a negative N that changes one visually verifiable binding. This P1/P2/N design evaluates whether a structural decision is stable across two valid expressions and also supplies two views of the same binding during training. An exploratory evaluation of pretrained checkpoints shows uneven transfer. Structure-CLIP remains strong on relation direction, but improves little or degrades on host, left–right, and depth bindings. We therefore introduce Complex-SCLIP, which applies the existing knowledge-enhanced encoder to both positives and jointly optimizes dual ranking, multi-positive contrastive learning, and positive consistency against the same structural negative. On the 240-instance test split, Complex-SCLIP reaches 69.2% Both-accuracy, outperforming Structure-CLIP (41.7%), five published compositional methods (32.1–39.6%), and an in-domain Structure-CLIP fine-tuning baseline trained with the released objective (48.3%). The improvement is consistent across three seeds and transfers to SUGARCREPE++. COCO 5K Mean Recall remains above CLIP (61.7% vs. 60.3%), though below the Structure-CLIP starting point.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.