Diffuse2Seg: Diffusion Models Can Segment Anything Without Supervision
Abstract
Open-world entity segmentation aims to predict masks for arbitrary visual entities across domains and at multiple granularities, from parts to whole objects. In this setting, SAM sets a strong standard: trained on SA-1B, comprising 11M images and over 1B carefully annotated masks, it achieves remarkable zero-shot performance. Collecting such labels is expensive and time-consuming, however, which limits how far this recipe can scale. Text-to-image diffusion models offer a way around this. Their intermediate features transfer well across perception tasks, and since object structure emerges as the model denoises a noise sample into an image conditioned on a text prompt, that structure is already encoded in these representations. It should therefore be possible to exploit them for open-world entity segmentation without retraining or supervision. To realize this, we present Diffuse2Seg, which repurposes generative diffusion models for automatic mask generation. This is achieved by propagating a grid of point prompts through their self-attention representations in an edge-preserving manner. Diffuse2Seg produces multi-granular instance masks and outperforms prior state-of-the-art label generators by 4.3-7.1 p.p. in across five domains. Training an instance segmentation model on these generated masks advances detector-free open-world segmentation by 7.4 and 7.7 p.p. on "things" and "stuff+things" datasets and surpasses the detector-based UnSAM on "stuff+things" by 2.1 p.p. in . Finally, we show that a model trained on Diffuse2Seg labels provides a strong initialization for semi-supervised learning, matching its fully supervised counterpart with only 5k labeled images and outperforming it with 10k. Code will be made publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.