SHAPE: Structural Hierarchy of Affinities for Patch Enhancement in Training-Free Open-Vocabulary Semantic Segmentation
Abstract
Training-free Open-Vocabulary Semantic Segmentation (OVSS) assigns text-specified class labels to image pixels without additional training. In this setting, frozen CLIP provides visual–text alignment, while a vision foundation model (VFM) supplies structural cues for localization. Despite this, predictions still suffer from patch noise, where isolated patches disagree with surrounding context, and region noise, where coherent object regions fragment. To address these issues, we formulate structural refinement as support selection: the VFM determines which patches contribute to each update, while CLIP remains the sole source of class evidence. We propose Structural Hierarchy of Affinities for Patch Enhancement (SHAPE), which introduces three complementary experts at different structural levels—Edge, Node, and Region—corresponding to relation, patch, and region support, respectively. The Edge Expert leverages VFM-derived patch affinities as relation-level support to aggregate CLIP visual features. The Node Expert reduces patch noise by refining each patch's class scores using its patch-specific VFM neighbors. Finally, the Region Expert reduces residual region noise by sharing class scores within coherent regions among patches whose predictions agree. Across eight benchmarks with a ViT-B/16 backbone, SHAPE achieves the best mIoU among weakly supervised and training-free methods on every dataset, reaching average scores of 44.0 and 47.9 without and with background classes, respectively. Code will be made publicly available upon publication.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.