acceptodds
Under review as a conference paper at ICLR 2027

Hyper-RIS: Hypergraph Consistency for Training-Free Referring Image Segmentation

Abstract

Training-free referring image segmentation (RIS) typically first generates class-agnostic mask candidates and then uses a pretrained vision–language model to select the candidate that best matches the referring expression. However, existing methods mainly score candidates independently or model only low-order interactions, making them prone to errors when multiple visually similar candidates compete for the same expression. We introduce Hyper-RIS, which improves candidate selection by explicitly modeling high-order consistency among candidate proposals. Our framework consists of two main components. First, Multi-Factor Scoring integrates complementary semantic cues from the subject, attributes, relations, and complete expression with dense spatial evidence to produce a unified score for each candidate. Second, Consistency Hypergraph Modeling uses the candidate scores and visual features to capture structured interactions among proposals. It first establishes pairwise visual context and then groups visually related candidates into hyperedges, where visually coherent groups provide consistency-weighted messages to their members. Candidate-wise reliability further suppresses conflicting group evidence, and the resulting high-order context is used to refine the proposal scores. Extensive experiments demonstrate that Hyper-RIS achieves state-of-the-art performance across multiple referring image segmentation benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.