acceptodds
Under review as a conference paper at ICLR 2027

SelRef: Selective Privileged Refinement for Multimodal Learning with Missing Modalities

Abstract

Multimodal learning has achieved remarkable success across a broad range of applications, yet its performance often deteriorates when one or more modalities are missing at inference. Fully observed multimodal data available during training can serve as privileged supervision. Existing distillation-based methods primarily transfer such knowledge by aligning predictions or representations between complete and incomplete inputs, leaving the usefulness of teacher guidance for each observed subset insufficiently explored. To tackle this limitation, we propose SelRef, a unified Selective Privileged Refinement framework that selectively exploits full-modality supervision to refine incomplete-modality predictions. SelRef first constructs a subset-conditioned prediction anchor by adaptively aggregating available modality experts learns an anchor-relative correction from their intermediate representations. During training, Privileged Task Gain and Current-State Compatibility jointly select teacher directions that improve upon the anchor and remain compatible with local descent at the refined prediction. We theoretically characterize the local descent properties of these directions and demonstrate the effectiveness of SelRef on three benchmarks spanning medical image segmentation and multimodal classification under complete and missing modality settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.