acceptodds
Under review as a conference paper at ICLR 2027

On Unidentified Representation Optimality of Vision-Language Compositional Reasoning

Abstract

Contrastive Language-Image Pre-training (CLIP) is a cornerstone of multimodal representation learning, while it is born with a persistent inability to reason about compositional relations. What's worse, a deficiency can be inherited by the Multimodal Large Language Models (MLLMs) that adopt its image encoder as their visual receiver. Why such failures arise has remained without a rigorous account. We develop a token-aware causal representation learning framework built on sequential, language-token Structural Causal Models (SCMs), recasting block identifiability for tokenized text and proving that CLIP's contrastive objective recovers the modality-invariant latent variables at the token level. It yields a principled identifiability-based diagnosis, i.e., the compositionality derived from a vanilla CLIP is nonidentifiable: pseudo-optimal text encoders attain perfect modality-invariant alignment while remaining provably invariant to the SWAP, REPLACE and ADD operators over atomic concepts, so no contrastive optimum guarantees separation between a caption and its hard negatives. Through the modality gap, the problem propagates to the image encoder, which leads to an algorithm-independent impossibility result: no projector, LLM capacity, or instruction tuning can lift a downstream MLLM above chance on the affected queries if they use the pseudo-optimal encoders to receive images. Guided by our theoretical recipe, we propose the remedy of hard negative strategies measurably repair both CLIP and MLLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.