The Wittgensteinian Representation Hypothesis: Is Language the Effective Reference Geometry of Multimodal Alignment?
Abstract
Understanding why independently trained neural networks from different modalities exhibit shared representational structures, and what governs this cross-modal alignment, remains an open question in representation learning. Existing evidence relies on symmetric similarity measures, which can detect modal alignment but are structurally blind to its direction. We introduce directional modal alignment analysis using cycle-kNN, an asymmetric alignment measure, and apply it to dozens of independently trained unimodal models spanning point clouds, vision, and language. We uncover a consistent directional asymmetry: non-language modalities align with the neighborhood structure of language significantly more than language aligns with theirs. This pattern holds across all model families and scales, yet remains entirely invisible to symmetric measures. Mechanistic analysis traces this directionality to feature-density asymmetry, whereby language representations occupy the most compact regions of representational space. The Information Bottleneck framework provides a principled interpretation: optimization under compression promotes alignment with the discrete, compositional structures characteristic of language. We formalize this insight as the Wittgensteinian Representation Hypothesis: the semantic structure of language constitutes the effective reference geometry for cross-modal representation alignment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.