acceptodds
Under review as a conference paper at ICLR 2027

Cross-modal Heterogeneous Privileged Policy Optimization for Medical Image Understanding

Abstract

Foundation models for medical imaging now span self-supervised vision, dense supervision, and vision–language learning, offering complementary knowledge about structure, localization, and semantics. Clinical deployment, however, still favors compact vision-only models, and transferring from such heterogeneous teachers jointly is difficult because their modalities, objectives, and representation spaces do not match. We introduce HiPPO, an on-policy distillation framework that transfers heterogeneous multimodal foundation models into a compact vision-only student. Conditioned on the current student state, HiPPO samples candidate representations as binary selections over a unified codebook, scores each in the native space of every applicable teacher via teacher-specific projectors and compatibility functions, and reinforces candidates with stronger group-relative feedback. The selected representations drive position-wise visual distillation from frozen visual foundation models and privileged semantic distillation from a frozen medical vision–language model conditioned on annotation-derived views available only during training. We further establish MedBridge-Bench, spanning case-level disease diagnosis and voxel-level lesion recognition over more than 18,000 volumetric CT, MRI, and PET/CT studies. Across these settings, HiPPO improves compact medical vision models and attains competitive performance, while all teachers, privileged inputs, and policy modules are discarded at inference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.