RFMedSAM 2: Automatic Prompt Refinement for Medical Image Segmentation with SAM 2
Abstract
The Segment Anything Model 2 (SAM 2) is a prompt-driven foundation model that generalizes SAM to both image and video segmentation, achieving strong zero-shot segmentation performance. However, despite its potential for medical image analysis, SAM 2 inherits key limitations from its predecessor, including binary mask outputs, inability to predict semantic labels, and reliance on precise prompts for target localization. Moreover, its direct application to medical image segmentation remains suboptimal. To investigate its upper performance bound, we fine-tune SAM 2 with targeted architectural modifications and oracle (ground-truth) prompts, reaching a Dice scores of 92.3% on BTCV, surpassing the state-of-the-art nnUNet by 9.8%; we stress that this is a diagnostic upper bound, not a comparable operating point, since it consumes ground-truth boxes and class labels. To eliminate the reliance on manual prompts, we develop an automatic prompt generator based on U-Net, which predicts coarse masks and bounding boxes that are subsequently refined through SAM 2’s progressive dual-stage inference. Furthermore, we propose QRefine-Former to transfer domain-specific medical knowledge from the U-Net prompt generator to the SAM 2 foundation model, further improving segmentation performance. Extensive experiments on the AMOS22, ACDC, and BTCV datasets demonstrate the superiority of our method, achieving Dice scores of 91.0% and 88.6% on AMOS22 Task 1 and Task 2, 93.6% Dice on ACDC, and 86.8% Dice on BTCV.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.