Deciding in the Dark: Multimodal LLMs as Evidence Selectors for Few-Shot Ultrasound Video Segmentation
Abstract
Few-shot segmentation of ultrasound sweeps depends on which annotated examples the segmenter sees and from which frame a mask is propagated, yet both are usually chosen by visual similarity, which is unreliable when speckle, shadowing and probe position change how the same anatomy looks. We present LucidSeg, a training-free framework in which a multimodal large language model (MLLM) makes discrete decisions about the evidence while frozen segmenters compute every mask. At Level 1, the MLLM chooses a small support set from retrieved candidates of other subjects, shown either as annotated exemplars or as annotations transferred onto the query, and a frozen in-context segmenter uses it; at Level 2, it chooses the starting frame of a sweep by judging the predicted outlines of candidate frames, and SAM2 propagates the mask from there. Because the MLLM never outputs a mask, box or point, its decisions serve any in-context segmenter. On four ultrasound datasets, a locally deployed open-weight MLLM matched a reranker trained on the target data and improved Dice over visual similarity by up to 0.11, with the largest gains where the target's appearance depends on anatomical side or probe position. On thyroid sweeps, the chosen starting frame came within 0.05 Dice of the best candidate; on spine sweeps over repeating vertebrae, starting from all candidates was better.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.