Protect What Is at Risk, Distill Where It Responds: On-Policy Distillation for Continual Learning in MLLMs
Abstract
On-policy distillation (OPD) can consolidate task experts for continual learning in multimodal large language models (MLLMs), but preserving experts does not ensure effective historical supervision. Changing visual evidence, instructions, and student responses expose historical capabilities unevenly, while the general OPD ignores which capabilities are being forgotten. We argue that **historical supervision must target vulnerable knowledge on trajectories that reveal it**. Guided by this principle, we propose **R**isk-guided **E**xpert **S**election via functional **P**robes for **ON**-policy **D**istillation (**RESPOND**), the first continual OPD framework to couple *what knowledge needs protection* with *where it can be distilled*. RESPOND probes current trajectories to supervise capabilities that are vulnerable in the student and observable through their experts, enabling consolidation with bounded historical-teacher queries per update. Experiments on eight sequential tasks and two OOD benchmarks show RESPOND improves final overall averages by 2.25–2.91 points over the strongest competing continual methods across three backbones. On Qwen3.5-4B, RESPOND reduces average forgetting by 50.2% and 69.7% relative to two vanilla continual OPD baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.