acceptodds
Under review as a conference paper at ICLR 2027

Zero-Shot Perception with Generator

Abstract

Pretrained on web-scale visual data, generation models acquire rich visual representations and expressive world priors that provide a natural foundation for perception. To unlock these capabilities, we explore perception as generation, leveraging generation paradigm to convert these internal representations directly into visual predictions. Although recent models such as Vision Banana demonstrate strong performance and impressive zero-shot generalization across diverse perception tasks, reliance on proprietary backbones hinders independent investigation and open reproduction. In this work, we present UniVision, an open 8B vision foundation model that unifies zero-shot visual perception and image generation. UniVision generates metric depth, surface normals, and semantic, instance, and referring segmentation as RGB images, which are converted back to target modalities via our noise-tolerant pixel codecs. UniVision achieves advanced zero-shot performance across all five tasks while capturing finer geometric detail than specialized baselines. We show that the core recipe behind these capabilities lies in deterministic, noise-tolerant codecs and scalable data pipelines. Furthermore, based on UniVision, we systematically investigate the perception-as-generation paradigm across both inference and training sides, revealing how inference-time generation techniques and test-time aggregation boost perception accuracy, characterizing task-specific data scaling dynamics, and demonstrating positive data synergy where joint perception training actively enhances generation quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.