Persistent Representational Misalignment Induces Vision-Language Model Failures
Abstract
Vision-language models (VLMs) fail at some tasks that are simple for humans—but why? We hypothesize that many such failures stem from VLMs’ difficulties translating between vision and language representations—a misalignment the Platonic Representation Hypothesis predicts should disappear with scale. This paper is composed of two parts: in the first, we show that existing vision and language representations are linearly unalignable. This misalignment is an inherent property of the data and hence expected to persist regardless of either scale or the alignment's complexity. In the second, we demonstrate how this misalignment both leads to a predictably induced generalization failure and correlates with existing failure modes, where more representationally aligned VLM backbones accordingly exhibit fewer failures. Overall, these results imply representations will not become aligned simply given sufficient scale, providing evidence against a strong version of the Platonic Representation Hypothesis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.