PartFactor: Recovering Unseen Parts with Decoupled Foreground Prediction
Abstract
Open-vocabulary part segmentation remains challenging: models must distinguish visually similar parts, recover fine spatial boundaries and transfer to categories absent from task supervision. Existing approaches strengthen object–part correspondence and spatial aggregation, but a correctly ranked part can still disappear when recognition uncertainty suppresses foreground. We propose PartFactor, whose object-conditioned decoder first learns conditional part identity by combining structural features, image–text matching and relations among parts. With this decoder fixed, a compact visual head then learns foreground support. A background-logit reparameterization makes foreground probability depend on visual evidence alone, allowing the head to recover missing regions while preserving the decoder’s conditional part distribution. Under the same full-image Pred-All evaluation, PartFactor improves unseen mIoU over a state-of-the-art baseline by 10.12 percentage points on PartImageNet and 23.79 points on PartImageNet-OOD. Fixed-decoder experiments trace the gains to recovery of correctly ranked parts, including at matched background false-positive rates. On PartImageNet, PartFactor needs only about 1/21 as many task-trainable parameters and achieves a 4.54× inference speedup over the same baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.