Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation
Abstract
Referring Expression Segmentation (RES) bridges vision and language but traditionally relies on extensive annotated datasets and struggles with implicit queries. While the Segment Anything Model 3 (SAM3) excels in promptable segmentation, its application to RES exposes **two critical challenges**: 1) SAM3 lacks the intrinsic capacity to decode lengthy, implicit expressions, and 2) it is highly vulnerable to both erroneous grounding priors and biased evaluation preferences from Multimodal Large Language Models (MLLMs), leaving its mask predictions unrefined. To address this, we introduce **Tarot-SAM3**, a novel training-free framework driven by the **key idea** of dual-stage refinement over MLLM-derived grounding cues and SAM3 mask predictions. Our **technical contribution** is two-fold. First, an Expression Reasoning Interpreter (**ERI**) is designed to guide the MLLM through structured parsing and evaluation-aware rephrasing, translating arbitrary queries into error-tolerant, heterogeneous prompts for SAM3 mask generation and performing consensus-driven mask selection within each prompt type. Second, the Mask Self-Refining (**MSR**) phase establishes a DINOv3-guided feedback loop, which identifies the most reliable mask and adaptively reconfigures point prompts to automatically resolve over- and under-segmentation. Extensive experiments demonstrate that Tarot-SAM3 achieves state-of-the-art results among MLLM-SAM3 architectures across MLLM scales, while under the Qwen2.5-VL 7B setting, it delivers **21.3%** and **28.2%** average relative improvements over the SAM3 Agent on explicit and implicit benchmarks, respectively. Open-world evaluations confirm our framework's superior generalization and robustness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.