MuseVLP: Clinically Grounded and Anatomically Consistent Vision-Language Pretraining for Multi-Sequence Brain MRI
Abstract
Brain MRI sequences depict the same anatomy through complementary contrasts, posing a central challenge for vision–language pretraining: grounding sequence-specific clinical evidence while learning sequence-invariant anatomical relationships. Existing coarse examination-level image–report alignment leaves both aspects unresolved, obscuring the sequence and anatomical sources of individual findings while providing no explicit supervision for shared anatomical relations. We introduce MuseVLP, a multi-sequence vision–language pretraining framework that is clinically grounded and anatomically consistent, addressing these challenges through multi-granular clinical grounding and relational consistency. MuseVLP decomposes examination-level reports into sequence-compatible descriptions and anatomically localized findings, replacing ambiguous global alignment with explicit alignment between clinical evidence and its supporting sequences and anatomical regions. Building on these grounded representations, we propose sequence-invariant anatomical relational consistency (ARC), which matches each region’s distribution of inter-region similarities to an exponential moving average reference aggregated from other sequences of the same examination, encouraging invariance in anatomical relations without requiring identical regional features. We pretrain MuseVLP on 87,675 MR-RATE examinations spanning T1-weighted, T2-weighted, FLAIR, and susceptibility-weighted MRI with paired radiology reports, and evaluate it across five brain MRI datasets for linear-probe classification, zero-shot classification, and report generation. MuseVLP consistently outperforms existing vision–language pretraining methods and transfers effectively to external datasets, suggesting that anatomical relations complement sequence-specific clinical grounding to learn transferable brain MRI representations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.