Referential Recalibration for Fine-Grained Referring Expression Segmentation
Abstract
Fine-grained referring expression segmentation extends classic RES from segmenting an entire referred object to localizing a specific part within a host-object. Most existing methods formulate this task as a direct mapping from a holistic expression to a target mask, which works well for object-level grounding. However, when the referred target is a fine-grained part of an object, the model is prone to assigning the referential focus to the host-object rather than the target-part. Moreover, even when the target-part is partially distinguished, the resulting score map may remain incomplete or drift into irrelevant regions. To address these issues, we introduce Referential Attribution Reformulation, which turns the host-object from a source of interference into useful guidance for target-part localization, and Bidirectional Collaborative Rectification, which further refines the initial target-part score map by recovering missing regions and suppressing drift-prone activations using internal target-part support and structural affinity. Our framework is training-free and improves VLM-based segmentation without additional fine-grained annotations or parameter updates. Extensive experiments demonstrate consistent improvements across different vision-language segmentation paradigms, supporting the effectiveness and broad compatibility of the proposed inference-time framework.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.