Beyond Multi-modal Alignment: An Information-Preserving Framework for Vision-Language Models
Abstract
Prevalent Vision-Language (VL) alignment techniques in Multi-modal Large Language Models (MLLMs) struggle to adequately align the language model with visual inputs, leading to hallucinations and undermining reliability. We rethink the modality alignment in MLLMs from the perspective of reducing information loss and present an efficient plug-in, VL Superior Alignment (VLSA), that decouples alignment into two stages. The first stage, Perception Alignment, couples compressive high-resolution encoding with reconstructive training through a pretrained text-to-image latent diffusion model (LDM), encouraging compact visual embeddings to preserve input content and align with language. The second stage, Cognition Alignment, introduces auxiliary codebook-index prediction, supplemented by pixel-value prediction, to train the LLM's understanding and use of visual information. Analyses of these two stages yield two empirical insights: reconstruction fidelity alone does not fully characterize the usefulness of compact visual representations for language tasks, and the LLM's utilization of visual information can improve even when the representations are held fixed. Extensive experiments across over 30 MLLM benchmarks and 9 MLLM architectures, thorough ablations, and analyses of computational overhead demonstrate that VLSA consistently improves performance over high-resolution baselines while maintaining inference costs close to those of low-resolution models, with further gains when the computational budget is increased. In service to the MLLM research community, our code and model checkpoints will be publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.