Context-Space Adaptation of Frozen Multimodal Large Language Models for Medical Image Diagnosis
Abstract
Multimodal large language models (MLLMs) offer a promising paradigm for medical image diagnosis, yet their performance can degrade substantially when applied to medical datasets with heterogeneous visual distributions and diagnostic tasks. Fine-tuning can improve domain adaptation, but requires repeated model updates for each target dataset. In this work, we investigate gradient-free adaptation of frozen MLLMs through the construction of dataset-specific context states. We propose Clinician Mimetic Workflow, a context-space adaptation framework that adapts frozen MLLMs by jointly constructing complementary visual and experiential context states. The framework comprises two coordinated context construction processes. Discriminative Exemplar Coreset Selection (DECS) organizes representative diagnostic cases in a frozen DINOv2 feature space and refines their adaptive retrieval keys to construct the visual context state. Self-Refined Experience Summarization (SRES) distills reusable diagnostic experience from diverse reasoning trajectories and maintains it through adaptive experience-state updates to construct the experiential context state. During inference, the visual exemplars provided by DECS and the diagnostic experience accumulated by SRES are jointly incorporated into the multimodal context, enabling the frozen MLLM to perform dataset-specific diagnostic reasoning without updating pretrained model parameters. We systematically evaluate the proposed framework on a subset of the public MedMNIST 2D datasets and a private nasopharyngeal carcinoma MRI dataset. Experimental results show that our method consistently improves over representative gradient-free in-context learning methods and achieves competitive diagnostic performance on real-world clinical data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.