LatentRefine-SegZero: Latent-Space Refinement for Referring Segmentation via Reflective Reasoning
Abstract
Significant progress has been made in referring segmentation by combining multimodal large language models (MLLMs) with the Segment Anything Model (SAM). However, balancing in-domain accuracy and out-of-domain generalization remains challenging. SFT-based methods achieve strong in-domain segmentation performance, but updating pretrained MLLM parameters with pixel-level losses may limit generalization to out-of-domain referring expressions. RL-based methods that avoid such updates typically control SAM through discrete box and point prompts, limiting fine-grained mask refinement. To mitigate this trade-off, we propose LatentRefine-SegZero, a latent-space refinement framework for referring segmentation based on reflective reinforcement learning. Our method comprises three components: (1) , which uses reflective reasoning to assess proposal reliability and select a refinement pathway; (2) , which converts proposal representations into supplementary latent prompts for mask refinement based on reliable proposals; and (3) , which converts reflective representations into new latent prompts for target relocalization. GRPO optimizes proposal generation and reliability assessment; pixel-level supervision trains lightweight adaptation modules without updating pretrained MLLM backbone weights. Experiments demonstrate that LatentRefine-SegZero achieves a favorable balance between in-domain segmentation accuracy and out-of-domain generalization. It leads the compared methods without pixel-loss updates to pretrained MLLM parameters by 1.4–3.9 cIoU points across all eight RefCOCO-series splits, while outperforming the compared methods with such updates on RefAdv and PR-Bench. It also improves over the Qwen2.5-VL+SAM2 baseline in both domains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.