acceptodds
Under review as a conference paper at ICLR 2027

PinPoint: Training-Free Referring Image Segmentation via Prompt Disambiguation

Abstract

Modern referring image segmentation pipelines couple a vision-language model (VLM) for grounding with a promptable segmenter such as the Segment Anything Model (SAM) for mask generation. Prior training-free instances of this recipe often trail fine-tuned and reinforcement-learning (RL)-tuned specialists, and it has been unclear whether the gap comes from the VLM's grounding, SAM's capacity, or the prompt. We show that prompt ambiguity is an important bottleneck in this architecture; a VLM-proposed bounding box (bbox) leaves SAM to guess which pixels inside the bbox belong to the object the expression denotes. Interior points are the natural disambiguator, but where they are placed matters; naively sampled points can land on boundaries, distractors, and background clutter, and even hurt performance compared to the bbox alone. Supervised and RL-tuned methods improve spatial prompts through task-specific training; we show that informative point prompts can instead be constructed without such training. At a matched budget of up to five interior points and on the same frozen stack, our selector improves cumulative Intersection-over-Union (cIoU) by 12–18 points across RefCOCO(+/g) splits. We introduce **PinPoint**, a deterministic, training-free point selector that fuses four visual cues into a consensus map, selects compact, spatially diverse points away from boundaries, and uses the frozen VLM to label each point and filter low-confidence labels. Without any task-specific training, PinPoint is competitive with supervised and RL-tuned specialists while issuing only two VLM calls per query.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.