Cross-modal Heterogeneous Privileged Policy Optimization for Medical Image Understanding
Abstract
Foundation models for medical imaging now span self-supervised vision, dense supervision, and vision–language learning, offering complementary knowledge about structure, localization, and semantics. Clinical deployment, however, still favors compact vision-only models, and transferring from such heterogeneous teachers jointly is difficult because their modalities, objectives, and representation spaces do not match. We introduce HiPPO, an on-policy distillation framework that transfers heterogeneous multimodal foundation models into a compact vision-only student. Conditioned on the current student state, HiPPO samples candidate representations as binary selections over a unified codebook, scores each in the native space of every applicable teacher via teacher-specific projectors and compatibility functions, and reinforces candidates with stronger group-relative feedback. The selected representations drive position-wise visual distillation from frozen visual foundation models and privileged semantic distillation from a frozen medical vision–language model conditioned on annotation-derived views available only during training. We further establish MedBridge-Bench, spanning case-level disease diagnosis and voxel-level lesion recognition over more than 18,000 volumetric CT, MRI, and PET/CT studies. Across these settings, HiPPO improves compact medical vision models and attains competitive performance, while all teachers, privileged inputs, and policy modules are discarded at inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.