acceptodds
Under review as a conference paper at ICLR 2027

From Global to Dense Representational Convergence: Shared Semantic Geometry Simplifies Vision-Language Connection

Abstract

Recent work on representational convergence suggests that independently trained vision and language models can develop similar semantic geometries despite differing modalities and training objectives. Existing studies, however, have mainly examined global image and sentence embeddings. In this paper, we ask whether this convergence also emerges in the dense patch representations of vision transformers. Using images with detailed pixel annotations, we capture how a broad range of concepts is organized in the dense vision space and compare this structure with the corresponding text representations. Interestingly, we find substantial alignment at this level as well, with consistently stronger alignment for larger and more recent vision and language models. Furthermore, we learn linear mappings between the vision and text representation spaces to exploit the consequences of this compatibility. Better-aligned spaces are easier to connect: their mappings are more accurate, require fewer training concepts, and perform better in downstream tasks. Specifically, stronger intrinsic alignment is associated with greater generalization to held-out categories in cross-modal retrieval, higher mIoU in text-guided semantic segmentation, and higher scores in zero-shot image captioning. Altogether, these findings extend evidence of representational convergence from global embeddings to spatially grounded concepts and link intrinsic alignment to easier cross-modal mapping and better downstream performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.