SPOR: Semantics-Preserving Ordinal Refinement for Open-Vocabulary Panoptic Segmentation
Abstract
Open-vocabulary panoptic segmentation requires recognizing categories unseen during training while accurately separating object instances and delineating coherent stuff regions. However, masks generated from appearance and semantic cues alone may erroneously merge semantically similar regions at different depth layers. We introduce SPOR (Semantics-Preserving Ordinal Refinement), a framework that uses within-image depth ordering to constrain spatial grouping in open-vocabulary panoptic segmentation. SPOR converts monocular depth estimates into normalized ordinal ranks, yielding a representation invariant to strictly increasing transformations of depth values. Rather than equating depth continuity with segment identity, it applies ordinal geometry selectively throughout mask formation. Ordinal affinities regulate local and cross-scale feature aggregation, while query-conditioned boundary routing uses reliable ordinal intervals to regulate cross-attention at uncertain boundaries while protecting confident mask cores. A training-only partition consistency objective combines ground-truth segment relations with ordinal gaps and cross-scale agreement to encourage within-segment coherence and cross-segment separation. Direct geometry-dependent modifications are confined to the mask-formation pathway, leaving pretrained semantic embeddings and the visual feature used for semantic pooling unchanged. Built on ODISE and fine-tuned on COCO Panoptic, SPOR achieves 55.5 PQ on COCO and 22.8 PQ in zero-shot transfer to ADE20K, while retaining broadly comparable performance across five open-vocabulary semantic segmentation benchmarks. These results suggest that ordinal scene structure can complement semantic recognition by constraining how pixels are grouped.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.