Partita: Multi-Granularity Promptable 3D Segmentation from Hierarchical Supervision
Abstract
Promptable 3D segmentation involves targets at different granularities, especially when proposing masks from a single click. However, single-granularity annotations limit candidate coverage and target diversity for interactive segmentation. We introduce Partita, a promptable 3D segmentation model that learns multi-granularity prediction and interactive refinement from hierarchical supervision. To provide this supervision at scale, we develop an automatic annotation pipeline that uses a vision-language model (VLM) to organize existing part annotations into nested hierarchies through recursive grouping and correction, yielding PartNeXt-XL, a dataset with 4.20M region masks across about 168K objects. Partita fuses low-resolution prompt-aware features with high-resolution geometric features to resolve finer parts and support efficient interactive refinement. Whereas prior promptable 3D methods supervise a single target mask per prompt, Partita uses multi-granularity matching to jointly supervise predictions across multiple levels. Experiments demonstrate improvements in interactive segmentation and multi-granularity prediction, as well as competitive automatic proposal generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.