mDINO: A sequential 2D–3D Latent-space Self-distillation Pre-training Framework for Medical Foundation Models
Abstract
Medical foundation models based on self-supervised learning (SSL) follow two paradigms: reconstruction-based methods capture the local texture at the expense of global semantics, and contrastive methods yield discriminative representations while losing fine-grained local details. Latent-space self-distillation offers a unified alternative to both paradigms, yet has not been adopted in medical foundation models, especially 3D domain, because of two main challenges: 1) Latent-space self-distillation methods demands large-scale pre-training corpora that far exceed the scale of available medical datasets. 2) Existing latent-space self-distillation models, such as DINOv3, are pre-trained for 2D natural images, leaving the domain gap of 3D medical foundation models unaddressed. In this paper, we propose mDINO, a sequential 2D-to-3D latent-space self-distillation pre-training framework for medical foundation models. To address data scarcity, we construct a large-scale CT corpus of 113M axial slices extracted from 1.13M volumes spanning diverse anatomical regions. mDINO first conducts 2D pre-training on 2D corpus, then leverages the spatial correspondence between 2D corpus and their source 3D volumes to transfer the learned representations as a structurally consistent initialization for 3D pre-training. To bridge the domain gap and enable latent-space self-distillation on 3D volumes, we further introduce 3D Hilbert flatten for locality-preserving token serialization and isotropy 3D-medical-RoPE for volumetric positional encoding. mDINO achieves superior performance over all reconstruction-based and contrastive-based medical foundation models on 40+ widely used 2D and 3D medical segmentation benchmarks. Code and pre-trained weights will be released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.