acceptodds
Under review as a conference paper at ICLR 2027

RESCU: RELATIONAL SCENE CONSTRAINT UNDERSTANDING FOR ZERO-SHOT REFERRING IMAGE SEGMENTATION

Abstract

Referring image segmentation asks a model to find the exact pixels an expression describes. Zero-shot methods almost universally reduce this to cross-modal retrieval: embed the expression and each candidate region in a shared CLIP space and take the best match. That formulation cannot represent relational constraintsleft_of, between, negation because no single embedding-space comparison ever sees the joint configuration of target and anchor. We introduce ReSCU (lational cene onstraint nderstanding), which builds an explicit relational scene graph over open-vocabulary proposals, parses the expression into a typed query, and evaluates it symbolically to produce a hypothesis - then arbitrates it against a strong hypothesis obtained by prompting a modern promptable-concept segmenter (SAM3) directly with the raw expression. The key idea, Dual-Hypothesis Mask Arbitration (DHMA), is to trust the appearance branch by default and promote the symbolic hypothesis only when a gate calibrated on held-out data says the parse is trustworthy. We study three such gates, of increasing capacity: hand-set thresholds, a sparse learned classifier, and a denser one over an expanded feature set. The hand-set gate () sits at statistical parity with matched SAM3-text and does not transfer across datasets. A sparse logistic gate over 17 features (), selected (grid fixed a priori, scored on held-out data, test touched once), produces a significant win on RefCOCOg (/ val/test mIoU, /pp, CI excluding zero) that transfers frozen to RefCOCO (/pp) and RefCOCO+ (/pp). Replacing lock C's classifier with a denser gradient-boosted one over 27 features () nearly doubles the margin (/ mIoU, /pp, CI test) and still transfers frozen to RefCOCO (/pp) and RefCOCO+ (/pp). Both gates lead every published zero-shot RIS number we could verify, including HybridGL (/) and RefChess (/), and both help all four expression types, not only the relational one lock A targeted. An oracle upper bound sits points above lock A's baseline; lock C closes of that gap and lock D , so roughly half remains unrealized under the best gate we found. Detection is not the bottleneck of gold targets survive to the proposal stagebut a frozen VLM used for spatial-edge prediction collapses to a single answer on held-out pairs, so our default edge backend stays geometric rather than learned. Lock D's larger gain also comes from intervening - more often through a harder-to-audit model, so we keep lock C as a sparse, interpretable alternative alongside lock D's headline result, and view closing the remaining oracle gap as the natural next step.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.