Vision as Unified Multimodal Generation
Abstract
We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed through the native text and image generation spaces of a unified multimodal model (UMM), without task-specific architectures. With this formulation, the single model **VisionUMM** matches leading task-specialized systems across structured visual understanding, dense geometric prediction, segmentation, and multi-view visual geometry. Natural-language instructions and optional visual prompts specify the task, target regions or views, and decoding convention. Responses are then generated as text for symbolic records, images for dense spatial targets, or mixed outputs for compositional tasks. To enable large-scale training, we convert heterogeneous computer vision annotations into instruction-response examples compatible with these native generation spaces. This conversion yields **VisionCorpus**, a computer-vision instruction-response corpus spanning text, image, and mixed text-and-image targets. Starting from an off-the-shelf pretrained UMM, VisionUMM is trained primarily on the VisionCorpus, using auxiliary multimodal data as a capability-preserving mixture and requiring no task-specific prediction heads or architectural changes. The resulting model covers detection, OCR, keypoints, segmentation, depth, surface normals, point maps, and camera pose estimation, and can follow language-defined variants that combine category, color, region, and other visual cues. These results suggest unified multimodal generation as a scalable route for integrating computer vision into general-purpose foundation models. The VisionUMM model and VisionCorpus are publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.