UniSpace: Unified Visual Representation and Scalable Multimodal Modeling
Abstract
Semantic vision encoders provide a powerful interface for multimodal understanding, but their final tokens lose fine-grained visual details needed for reconstruction and image editing. We ask whether understanding, generation, and editing can share a visual representation built from a pretrained semantic ViT. A controlled probe shows that replacing only the patch embedding improves recoverability while keeping all Transformer blocks frozen. Motivated by this observation, we introduce Patch Reparameterization: the original semantic pathway is preserved, and a reconstruction-aware patch embedding supplies visual details to the same frozen backbone. We compress and concatenate the two token streams and balance their contributions during flow-matching training. The resulting tokenizers retain aggregate multimodal understanding while enabling high-fidelity reconstruction and image generation. PR-DINOv2 achieves an ImageNet reconstruction FID of 0.14 and a generation FID of 1.87 with classifier-free guidance. We further scale this representation into UniSpace, an 8B Mixture-of-Transformer-Experts model that supports understanding, generation, and editing without a separate VAE pathway. It scores 4.28 on ImgEdit and 0.84 on GenEval, demonstrating a shared ViT-based visual interface at system scale.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.