acceptodds
Under review as a conference paper at ICLR 2027

Probing Fine-Grained Semantic Alignment Across Vision and Language with Minimal Pairs

Abstract

To what extent does alignment between vision and language preserve specific semantic distinctions? We introduce a new benchmark that examines this question across shallow mappings between independently trained vision and language models, CLIP-style contrastive models, and generative vision-language models. Each benchmark item leverages a minimal pair (of pairs) design comprising two minimally contrasting images and two corresponding captions, with contrasts spanning object properties and relations, physical events, and social interactions. By testing whether similarity scores or matching decisions consistently favor the correct image–caption correspondences over the mismatched alternatives, the benchmark probes whether alignment preserves the specific semantic distinction targeted by each item. Across these settings, we find that alignment is highly conditional. Shallow unimodal alignment recovers some shared object and scene structure, but performs poorly on minimal pairs that require preserving specific relational, causal, or social distinctions. Multimodal training improves performance, but gaps remain even for stronger models compared to human baselines. Further analyses show that implicit captions are consistently harder than explicit captions, suggesting that current models are better at matching named visual properties than inferring latent consequences from world knowledge. Together, our results suggest that broad semantic structure may support cross-modal matching, but fine-grained inferential distinctions needed to match closely related images and captions remain fragile even in multimodal systems.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.