Negation-Aware Reranking Framework for Zero-Shot Referring Image Segmentation
Abstract
Zero-shot referring image segmentation reduces the reliance on task-specific expression–mask annotations by using pretrained vision–language models and class-agnostic mask proposal networks. However, existing methods show performance degradation on expressions containing negation, as CLIP-based holistic matching may select a candidate with high target compatibility despite violating the negation constraint imposed by the referring expression. To address this challenge, we introduce NARF, a training-free negation-aware framework that reranks candidate masks using a base score that preserves the base method's original assessment, a target-compatibility score that measures compatibility with the intended entity, and a negation-violation score that penalizes detected violations. We further construct a Negation-Focused Evaluation Set from existing RIS datasets for systematic evaluation. Across five zero-shot RIS methods, NARF improves aggregate oIoU and mIoU and provides further gains when combined with NegationCLIP, demonstrating the complementarity of negation-aware representation enhancement and candidate-level constraint verification.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.