Let Vision-Language Models Generate Images through Their Representations
Abstract
Recent unified multimodal models (UMMs) aim to combine visual understanding and generation within a single model, yet effectively integrating the two capabilities remains challenging. Leveraging pretrained vision-language models (VLMs) offers a way to build on already learned multimodal knowledge, but making this knowledge effectively accessible to visual generation remains nontrivial. We find that a central challenge lies in the heterogeneous visual representations used for understanding and generation. To address this challenge, we introduce Integral, which converts a pretrained VLM into a UMM by constructing a shared visual representation space across the two capabilities. Integral implements this design through two complementary components. Abstraction Fusion constructs shared visual tokens by aggregating hierarchical representations from the frozen vision encoder, while Block-wise MoT introduces a lightweight generative stream that interacts with the VLM only at block boundaries. Our approach preserves nearly all of the pretrained VLM's understanding capability while extending it to strong visual generation. Controlled comparisons further show that sharing the VLM representation space improves generative learning beyond what can be explained by reconstruction quality alone. Our results indicate that pretrained VLM representations already contain knowledge useful for visual generation, and that preserving their representation hierarchy enables this knowledge to be used without sacrificing visual understanding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.