Partition Everything: Whole-Image Segmentation by Palette-Conditioned Flow Matching
Abstract
Decomposing an image into its constituent objects and parts provides a fundamental representation for scene understanding and image editing. Yet whole-image segmentation remains cumbersome with promptable models such as SAM, requiring dense point prompting followed by heuristic merging and suppression of overlapping, duplicate masks. We introduce Partition Everything Model (PEM), which formulates whole-image partitioning as palette-conditioned image generation. Our rectified-flow model learns to repaint each region with the mean color of a supplied continuous palette field, representing the partition as an image without prescribing semantic labels or the number of regions. To distinguish regions at different spatial scales, we condition the model on two complementary palettes and jointly predict their corresponding color maps, combining broad spatial variation with local color contrast. Together, these maps provide a six-dimensional representation from which we extract an exhaustive, non-overlapping partition. Furthermore, we introduce annotated-region conditioning to learn from incomplete annotations while predicting complete image partitions at inference. Across seven benchmarks containing 12,138 images and 117,300 instances, PEM outperforms SAM3 by 15.57 percentage points in instance-weighted mIoU (55.27% vs. 39.70%), while predicting all regions jointly in a single forward pass.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.