Uncovering Modality Discrepancy: Rethinking the Validation of Foundation Models for General-Purpose 3D Medical Segmentation
Abstract
Foundation models have emerged as a transformative paradigm in 3D medical imaging, with the promise of unified quantitative analysis across diverse targets and imaging modalities. However, these models are predominantly developed and evaluated on datasets largely concentrated around a limited set of imaging modalities and anatomical regions. To provide an evaluation of out-of-the-box robustness of these foundation models, we curate a large-scale dataset comprising 490 whole-body PET/CT and 464 whole-body PET/MRI scans (approximately 675k 2D images and 12k 3D annotations) and conduct a thorough and comprehensive evaluation of representative 3D segmentation foundation models. Our analysis reveals a substantial gap between benchmark-reported performance and real-world generalization, with marked degradation on previously unseen data distributions and particularly severe failures on functional imaging modalities. These findings suggest that current foundation models remain far from achieving true universality. We suggest a reconsideration of how universality is defined and validated, extending evaluation beyond regional structural benchmarks toward whole-body structural and functional imaging. Bridging this gap will be essential for translating foundation models from controlled evaluation settings to real-world medical practice.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.