Prompt Variation Alignment for Robust Segmentation
Abstract
Promptable segmentation models such as the Segment Anything Model (SAM) family provide convenient interfaces for interactive mask generation through clicks, bounding boxes, text prompts, or their combinations. However, SAM-based segmentation models are often sensitive to the user prompt: clicks placed on different parts of an object, or different text prompts for the same object, can substantially degrade the predicted mask — making the model unreliable in practice, especially in critical applications such as medical imaging. To address this, we propose Best-Teaches-Rest (BTR), a post-training objective that enforces the alignment of predictions across prompts that refer to the same target object. BTR groups prompts that refer to the same object, identifies the highest-IoU prediction as the teacher, and trains the remaining predictions to match it — with only the teacher receiving ground-truth supervision. Unlike standard supervised fine-tuning where prompt invariance can only emerge implicitly across independent prompt–mask pairs, BTR enforces it directly through the loss. BTR loss can be applied uniformly to click variation, text variation, and joint click–text variation. We post-train SAM 3 with BTR on EntitySeg and observe consistent gains in both mask quality and prompt consistency over pretrained SAM 3. These gains transfer zero-shot to six object-centric instance segmentation datasets and to the out-of-distribution RAOS benchmark luo2024raos for medical image segmentation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.