acceptodds
Under review as a conference paper at ICLR 2027

Let Vision-Language Models Generate Images through Their Representations

Abstract

Recent unified multimodal models (UMMs) aim to combine visual understanding and generation within a single model, yet effectively integrating the two capabilities remains challenging. Leveraging pretrained vision-language models (VLMs) offers a way to build on already learned multimodal knowledge, but making this knowledge effectively accessible to visual generation remains nontrivial. We find that a central challenge lies in the heterogeneous visual representations used for understanding and generation. To address this challenge, we introduce Integral, which converts a pretrained VLM into a UMM by constructing a shared visual representation space across the two capabilities. Integral implements this design through two complementary components. Abstraction Fusion constructs shared visual tokens by aggregating hierarchical representations from the frozen vision encoder, while Block-wise MoT introduces a lightweight generative stream that interacts with the VLM only at block boundaries. Our approach preserves nearly all of the pretrained VLM's understanding capability while extending it to strong visual generation. Controlled comparisons further show that sharing the VLM representation space improves generative learning beyond what can be explained by reconstruction quality alone. Our results indicate that pretrained VLM representations already contain knowledge useful for visual generation, and that preserving their representation hierarchy enables this knowledge to be used without sacrificing visual understanding.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.