LLM-Guided Contrastive Learning for Cross-Modal Medical Image Segmentation
Abstract
Multimodal information has shown great potential in medical image segmentation, particularly when textual semantic cues are effectively exploited. However, existing cross-modal medical image segmentation methods often suffer from coarse cross-modal fusion and alignment due to two factors: (i) a lack of challenging negative examples for contrastive optimization, and (ii) the neglect of spatial correspondences between textual semantics and local image regions. To this end, we propose a cross-modal medical image segmentation framework via large language model (LLM)-guided contrastive learning (LGCSeg), which enhances representation learning and modality alignment. LGCSeg leverages a multimodal LLM to generate textual descriptions of potential infection regions in input images, which serve as semantic references for identifying relevant hard negative samples from a domain-specific medical text corpus. These samples drive contrastive learning, enhancing the discriminative power of visual-language representations. Furthermore, a multidimensional cross-modal interaction block is introduced to enable semantic interaction between visual and textual features along global, horizontal and vertical spatial dimensions. This design achieves both global and spatial-level alignment, facilitating accurate localization of text semantics within corresponding image regions. Experimental results demonstrate that LGCSeg outperforms the base model by 1.95% in Dice and 2.83% in mIoU on QaTa-COV19 (X-ray), and by 13.04% and 19.44% on MosMedData+ (CT), respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.