VITAL: Visual and Task Advantage Learning for On Policy Distillation in Medical Diagnosis
Abstract
Medical vision-language models (VLMs) have shown strong diagnostic performance. On-policy distillation (OPD) transfers these capabilities to smaller VLMs for faster inference and easier deployment. However, existing OPD methods often overlook what the student actually lacks. We uncover two counter-intuitive forms of student deficit: (1) Visual Deficit: misaligned rather than simply weaker visual dependence reduces sensitivity to clinical findings, causing diagnostic errors. (2) Task Deficit, where student rollouts are not consistently inferior to the teacher, and indiscriminate imitation overwrites useful student behaviors and impair generalization. Motivated by these observations, we propose VITAL, a VIsual and Task-Advantage Learning framework for on-policy distillation in medical diagnosis. VITAL improves distillation from two complementary perspectives. Specifically, Visual Advantage Policy optimization (VAP) estimates the teacher’s token-level visual dependence through image perturbations and scales distillation accordingly to improve the student’s visual dependence. Task Advantage Distillation (TAD) estimates the teacher's task advantage over each student rollout and rejects rollouts without positive teacher advantage, avoiding redundant or potentially harmful imitation. Experiments across four adult and neonatal datasets show VITAL outperforms prior OPD methods in-domain, while boosting out-of-domain F1 by 4.60% (CheXpert Plus) and 2.79% (NeoCXR-EV) over the strongest baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.