acceptodds
Under review as a conference paper at ICLR 2027

TAG: Calibrating SAM3 for Universal Image Segmentation without Dense Supervision

Abstract

The recent Segment Anything Model 3 (SAM3) extends powerful promptable segmentation from spatial cues to concepts, raising the prospect of a single foundation model for universal image segmentation. However, this promise often fails to translate into reliable performance on downstream segmentation tasks. In this work, we empirically identify concept recognition, rather than mask quality, as the primary bottleneck: SAM3 accurately delineates the objects it recognizes yet struggles to robustly determine which concepts are present. Driven by this, we present Text Anchor re-Grounding (TAG), a simple yet effective framework that rectifies the concept recognition of SAM3 on the target domain without dense annotations. Specifically, TAG introduces an Anchor Residual Pathway comprising image-conditioned, low-rank adapters to re-ground text anchors in the target distribution and modulate subsequent text-conditioned interactions. Guided solely by image-level category labels, TAG jointly optimizes binary cross-entropy with Soft Boundary Ranking objectives to calibrate presence confidence and distinguish present concepts from confusable negatives. Building upon this, we devise Confidence-Weighted Distillation to couple rich semantic context of a multimodal large language model with fine-grained spatial priors of SAM3, thereby achieving effective calibration in the absence of manual annotations. Extensive experiments demonstrate that TAG unlocks SAM3's segmentation potential with only 3.1M trainable parameters, achieving impressive performance across diverse scenarios, and notably, even surpassing leading mask-supervised methods on several metrics. Our code will be publicly available. Anonymous github link: https://anonymous.4open.science/r/TAG-9038.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.