Knowledge-Guided On-Policy Self Distillation for Medical MLLMs Continual Learning
Abstract
Medical multimodal large language models (MLLMs) need continuous adaptation to emerging clinical tasks while preserving previously acquired capabilities, yet collecting high-quality trajectories and designing explicit rewards remain difficult. On-policy self distillation (OPSD) enables continual learning by supervising the model on trajectories generated via its own evolving policy without complex reward design, but the effectiveness depends on the quality of teacher-side privileged signal. In this work, we observe that OPSD-based medical MLLMs continual learning is not simply to provide more medical context, but to construct privileged knowledge conditions that yield reliable teacher supervision on uncertain student-generated trajectories. With this insight, we propose knowledge-guided on-policy self distillation for medical MLLMs continual learning (MedKgSD), a reward-free and external-teacher-free framework that uses biomedical knowledge graphs to construct training-time privileged knowledge conditions for on-policy self-distillation. For each new task, MedKgSD builds a knowledge graph induced task-level concept space and derives sample-specific privileged conditions from biomedical knowledge relations to interpret student-visited states under structured medical context. Since knowledge-derived conditions could be noisy or mismatched, we introduce token-selective entropy calibration to identify reliable teacher signals on informative student states and distills the teacher behavior along student-generated trajectories. This converts structured medical relations into dense on-policy supervision for new-task adaptation, with prior-policy distillation reducing the forgetting of previous capabilities. Experiments across multiple medical backbones demonstrate that MedKgSD improves the adaptation-retention trade-off in continual medical adaptation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.