HiUni: Harmonizing Understanding and Generation in Shared Visual Space
Abstract
Unified multimodal models (UMMs) aim to support visual understanding and generation within a single framework. However, prevailing approaches often separate their visual representations or backbone computation to accommodate the different requirements of the two tasks, which limits the degree of unification. In this paper, we present HiUni, which revisits this design choice and explores whether a pretrained VLM can be directly extended into a unified model with a high degree of sharing. Specifically, HiUni reuses the same language aligned ViT and dense Transformer backbone for both capabilities, while confining task specific computation to the output side. Yet such a design introduces two key challenges. Direct generation in pixel space is difficult to optimize, while visual compression and semantic abstraction in pretrained VLMs may discard fine spatial details required for synthesis. To address these challenges, HiUni introduces a Pixel Diffusion Head tailored to direct pixel-space optimization and a lightweight HiGate pathway that provides fine-grained spatial details near the output stage. Extensive experiments show that HiUni achieves a similar average understanding score to the pretrained VLM while attaining competitive performance on image generation and editing. Controlled comparisons further demonstrate that a single dense backbone can effectively support both capabilities without task-specific backbone computation, highlighting a simpler path toward unified multimodal modeling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.