Erase-to-Verify: Agentic Self-Correction for Referring Image Segmentation
Abstract
Referring image segmentation aims to identify the object described by a natural-language expression and predict its pixel-level mask. Although frozen vision-language models and promptable segmenters provide a practical training-free solution, their predictions may select the wrong instance or cover only part of the intended target. We propose Erase-to-Verify, a training-free framework that treats each predicted mask as a hypothesis to be verified. Given a candidate mask, we locally blur its region and ask whether the referred target remains visible in the resulting image. Visible residual evidence indicates that the current mask is insufficient and should be refined, whereas target absence provides an operational stopping signal. By intervening on the image with the predicted mask, the framework obtains new visual evidence instead of repeatedly judging the prediction within the original observation. For a rejected mask, the framework further diagnoses the failure type. If the prediction corresponds to a wrong instance, it re-localizes the target from the target-suppressed image. If the prediction covers the correct target only partially, it zooms into the residual region and generates point prompts for SAM3 to recover the missing parts. The updated mask is then verified again within the same loop. All components remain frozen, requiring no task-specific fine-tuning, agent training, learned rewards, or learned routing policies. Experiments on all eight RefCOCO-family splits show that our method improves over the strongest reported training-free baseline in our comparison by 6.0 gIoU points on average. Results on ReasonSeg further demonstrate its generalization to implicit reasoning expressions. Our code will be made publicly available soon.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.