Refining Plane Priors for Query-Conditioned Fetal Ultrasound Segmentation
Abstract
Text-guided fetal ultrasound segmentation requires locating the structure specified by a clinical query among multiple structures in the same image. Existing multimodal large language model (MLLM) approaches rely primarily on internal representations that lack fetal-ultrasound domain knowledge. A pretrained fetal-ultrasound vision-language model can provide domain-specific visual features and plane-level activation maps. However, the visual features do not directly provide an explicit spatial prior, while the activation maps remain coarse and unstable for query-specific localization. We introduce PRISM (Plane-prior Refinement and Integration for Segmentation with MLLMs), a framework that converts plane-level evidence into query-conditioned spatial prompts. PRISM uses the pretrained model through two complementary pathways: its domain-specific visual features are supplied to the MLLM, while its plane activation maps provide spatial evidence for refinement. Activation Prior Refinement (APR) learns a residual correction of the activation map from visual context and query semantics, producing a target-specific spatial prior. Patch-Adaptive Prior Composition (PAPC) then constructs experts from the refined prior, visual features, and query representation, selects an expert subset with a Set Router, and fuses their predictions spatially with a Patch Router. Extensive experiments on public and private fetal ultrasound benchmarks demonstrate that PRISM consistently outperforms the compared methods across model scales, reaching 80.55% and 71.41% cIoU on the public and private benchmarks, respectively, with a 13B language model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.