Language Shapes Vision: Bidirectional Information Flow in Multimodal LLMs
Abstract
MLLMs have shown strong capabilities in jointly processing visual and linguistic information, making the underlying mechanism of multimodal integration a fundamental question. However, existing studies mainly examine autoregressive (AR) MLLMs, where causal attention has largely confined our understanding of this mechanism to a unidirectional flow from vision to language. In this work, we revisit multimodal integration mechanisms by studying diffusion MLLMs (DMLLMs) with full attention between vision and language. Surprisingly, we find that DMLLMs coordinate bidirectional information flow in a distinctly staged manner: language-to-vision flow dominates in shallow layers, and cooperates with vision-to-language flow in deeper layers. We further reveal that this language-to-vision flow prompts the alignment of question-relevant visual states with semantics and facilitates subsequent vision-to-language transfer. Motivated by this finding, we introduce language-to-vision flow into the training of conventional AR MLLMs and find that it brings more question-targeted processing of visual information, improving vision-centric understanding ability. Our findings reveal bidirectional information flow as an important mechanism underlying multimodal integration and provide insights for designing next-generation MLLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.