Revisiting Multimodal Pre-training for Vision-Language Alignment in MLLMs
Abstract
Vision-language alignment is a key prerequisite for MLLMs to perceive and understand the visual world. Although existing MLLMs typically employ a multimodal pre-training stage for building alignment, we find that the fine-grained, patch-level cross-modal correspondence remains insufficiently established. As text-to-image (T2I) attention plays a critical role in building cross-modal correspondence, we examine it on a pre-trained model and find that T2I attention rapidly shifts from broadly distributing over the image to concentrating on a small subset of textually irrelevant visual tokens, termed visual attention sinks, rather than attending to text-relevant visual evidence. Through further analyses of visual sinks' attention flow and training dynamics, we observe that they emerge in the early loss-reduction phase and persistently hinder the projector from learning fine-grained cross-modal correspondence. To address this issue, we propose Visual Sink-aware Alignment (ViSAlign), a simple and effective method that temporarily blocks the visual sinks' attention flow during multimodal pre-training. By preventing visual sinks from dominating the attention, it encourages the model to capture text-relevant visual evidence, thus learning a finer-grained cross-modal correspondence. Extensive experiments across diverse multimodal settings demonstrate that ViSAlign improves projector learning, strengthens vision-language alignment, and yields broad improvements over a variety of downstream tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.