TIGR-Seg: Text-Initialized Gaussian Refinement with Spatially Aware Visual Adaptation for Medical Image Segmentation
Abstract
Medical image segmentation remains challenging due to the limited availability of pixel-level annotations, heterogeneous target appearance, and variations across imaging modalities. Pretrained vision encoders provide strong representations, but updating all backbone parameters introduces substantial trainable capacity, while textual information often provides semantic and coarse spatial guidance rather than precise pixel-level localization. This work presents TIGR-Seg, a parameter-efficient text-guided segmentation framework comprising Gaussian-Gated Feature Adaptation (GGFA) and a Text-Initialized Gaussian Refinement (TIGR) Decoder. GGFA adapts a frozen pretrained vision encoder through context-conditioned feature modulation and a dynamically predicted correlated Gaussian spatial gate, enabling adaptation across both feature and spatial dimensions while updating only a small set of parameters. The TIGR Decoder incorporates stage-specific textual guidance across multiple decoding resolutions. At each stage, text initializes a Gaussian spatial hypothesis that is subsequently refined using visual evidence through reliability-weighted corrections. The refined spatial priors then guide feature aggregation, reconstruction, and progressive decoder refinement. Experiments across gastrointestinal endoscopy (BKAI), brain MRI (BTMRI), and breast ultrasound (BUSI) datasets demonstrate strong segmentation performance with a low trainable-parameter count. Additional evaluations support the effectiveness and generalizability of TIGR-Seg.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.