Modality Gap–Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models
Abstract
Despite the success of multimodal contrastive learning in aligning visual and linguistic representations, a persistent geometric anomaly, the Modality Gap, remains: embeddings of distinct modalities expressing identical semantics occupy systematically offset regions. Prior approaches to bridge this gap are limited by oversimplified isotropic assumptions, hindering their application in large-scale scenarios. In this paper, we address these limitations by characterizing the geometric shape of the modality gap and leveraging it for efficient model scaling. We reveal the structured geometric characteristics of the Modality Gap: it not only contains a global centroid shift, but also exhibits significant anisotropic residuals, while the two modalities still share similar dominant geometric structures. Guided by this precise modeling, we introduce ReAlign, a training-free modality alignment strategy. Utilizing statistics from massive unpaired data, ReAlign aligns text representation into the image representation distribution via Anchor, Trace, and Centroid Alignment, thereby explicitly rectifying geometric misalignment. Building on ReAlign, we propose ReVision, a scalable pre-training paradigm for Multimodal Large Language Models (MLLMs). ReVision integrates ReAlign into the pretraining stage, enabling the model to learn the distribution of visual representations from unpaired text before visual instruction tuning, without large-scale, high-quality image-text pairs. Our framework demonstrates that statistically aligned unpaired data can effectively substitute for expensive image-text pairs, offering a robust path for the efficient scaling of MLLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.