S3-Adapt: Grounded Semantic Transfer between Foundation Models for Medical Image Segmentation
Abstract
In medical segmentation cascades built around independently encoded image-text similarity and decoder prompts, language can guide mask prediction without shaping the segmentation encoder's spatial representations. We introduce S3-Adapt, a framework that connects a pre-aligned vision-language model with a promptable segmentation foundation model through Spatial Grounding, Semantic Transfer, and Selective Correction. Reciprocal Semantic Grounding integrates query-local visual context and reciprocal region-token associations into the source visual encoder, with dense grounding supervision shaping query-conditioned spatial features. Hierarchical Cross-Foundation Semantic Transfer projects these features into multiple stages of the segmentation encoder and jointly conditions gated residual updates on source semantics and the current target state. Point, continuous-mask, and text prompts complement this representation-level transfer to form a reference prediction. Evidence-Conditioned Selective Correction then combines local image evidence with grounding and prediction context to produce a single signed correction gate. The gate controls the direction and bounded magnitude of a logit update while the reference pathway remains fixed. The original parameters of both pretrained encoder backbones remain frozen. Experiments on four public medical image segmentation datasets demonstrate the effectiveness of S3-Adapt and competitive segmentation performance with only 1.97M task-adapted parameters. Source code is provided in the supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.