On Unidentified Representation Optimality of Vision-Language Compositional Reasoning
Abstract
Contrastive Language-Image Pre-training (CLIP) is a cornerstone of multimodal representation learning, while it is born with a persistent inability to reason about compositional relations. What's worse, a deficiency can be inherited by the Multimodal Large Language Models (MLLMs) that adopt its image encoder as their visual receiver. Why such failures arise has remained without a rigorous account. We develop a token-aware causal representation learning framework built on sequential, language-token Structural Causal Models (SCMs), recasting block identifiability for tokenized text and proving that CLIP's contrastive objective recovers the modality-invariant latent variables at the token level. It yields a principled identifiability-based diagnosis, i.e., the compositionality derived from a vanilla CLIP is nonidentifiable: pseudo-optimal text encoders attain perfect modality-invariant alignment while remaining provably invariant to the SWAP, REPLACE and ADD operators over atomic concepts, so no contrastive optimum guarantees separation between a caption and its hard negatives. Through the modality gap, the problem propagates to the image encoder, which leads to an algorithm-independent impossibility result: no projector, LLM capacity, or instruction tuning can lift a downstream MLLM above chance on the affected queries if they use the pseudo-optimal encoders to receive images. Guided by our theoretical recipe, we propose the remedy of hard negative strategies measurably repair both CLIP and MLLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.