MediVerse: A Unified Medical Vision–Language Foundation Model with Zero-Forgetting Continual Modality Expansion
Abstract
A unified medical foundation model should support many imaging modalities, provide a language interface, and expand without losing existing capabilities. Current models usually satisfy only part of this goal. Vision-only models support several modalities but cannot perform text-prompted zero-shot classification or image–text retrieval. Vision–language models provide these abilities but often require large paired corpora and may forget earlier modalities when adapted to new ones. We introduce MediVerse, a pair-efficient and extensible medical vision– language model. MediVerse uses a Multiway Transformer with a Domain-Aware Mixture-of-Experts. A deterministic router sends each image to its modalityspecific expert, while a separate text pathway interacts with the image pathway through contrastive alignment. We first pretrain on 19.4M unpaired images and then align the model using 3.2M image–text pairs. With a ViT-Large backbone, MediVerse wins 10 of 11 zero-shot benchmarks (mean 0.674 vs. 0.594 for the strongest baseline), is best or tied-best on all 10 linear probes (0.856 vs. 0.824), and wins all 4 cross-modal retrieval benchmarks (20.70 vs. 12.09) across nine imaging modalities. To add a new modality, we train a new routed expert while freezing the existing computation paths. This design preserves old-modality embeddings when the existing computation paths and inference conditions remain unchanged. Across fundus and OCT expansion, adapted baselines lose 2.97–6.11 points of pathology retrieval, whereas MediVerse shows no change on the evaluated pathology task.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.