acceptodds
Under review as a conference paper at ICLR 2027

Persistent Representational Misalignment Induces Vision-Language Model Failures

Abstract

Vision-language models (VLMs) fail at some tasks that are simple for humans—but why? We hypothesize that many such failures stem from VLMs’ difficulties translating between vision and language representations—a misalignment the Platonic Representation Hypothesis predicts should disappear with scale. This paper is composed of two parts: in the first, we show that existing vision and language representations are linearly unalignable. This misalignment is an inherent property of the data and hence expected to persist regardless of either scale or the alignment's complexity. In the second, we demonstrate how this misalignment both leads to a predictably induced generalization failure and correlates with existing failure modes, where more representationally aligned VLM backbones accordingly exhibit fewer failures. Overall, these results imply representations will not become aligned simply given sufficient scale, providing evidence against a strong version of the Platonic Representation Hypothesis.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.