SEmO: Seg+Edge minus Others Token with Progressive Mask Decoding for Reasoning Segmentation
Abstract
Reasoning segmentation aims to generate precise segmentation masks from complex and implicit language queries. Prior approaches mainly rely on a single <SEG> token to represent all segmentation-relevant information and predict the final mask in a single decoding stage. This design makes it difficult to capture fine-grained target edges while suppressing false positives in the background. We propose SEmO, a novel framework that augments the conventional <SEG> token with two role-specific tokens, <EDGE> and <OTHERS>, together with a Progressive Mask Decoder(PMD). The three tokens explicitly represent the target region, its edge, and the background, respectively, providing complementary information for mask prediction. At each resolution stage, PMD aggregates their role-specific predictions and propagates the refined result to the next stage for coarse-to-fine mask refinement. These designs enable SEmO to predict more accurate target edges while effectively suppressing background false positives. Experiments demonstrate that SEmO achieves state-of-the-art performance on ReasonSeg and improvements on referring segmentation benchmarks. Moreover, PMD requires only 0.69M parameters and achieves faster inference than existing mask decoders, showing that the performance gains are obtained without sacrificing efficiency. All code and model parameters of SEmO will be made available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.