LAMP: Learnability- and Horizon-Aware On-Policy Distillation from Privilege Modality
Abstract
Multimodal large language models (MLLMs) often struggle with perceptual information that is difficult to recover from RGB alone. Cross-modal on-policy distillation (OPD) addresses this issue by allowing the teacher to access privilege modalities such as depth or video, but existing methods treat token-level supervision uniformly, conflating transferable reasoning with unlearnable sensor information. Therefore, we propose LAMP, a learnability- and horizon-aware privilege-modality on-policy distillation framework for MLLMs. LAMP makes a two-fold contribution: First, it performs counterfactual modality attribution to identify whether each token’s teacher advantage is grounded in the student or privilege modality, enabling a self-calibrating gate to suppress untransferable supervision. Second, it introduces horizon-conditioned divergence weighting to emphasize tokens that have greater downstream influence on the reasoning trajectory. Extensive experiments on depth-to-RGB and video-to-RGB transfer show that LAMP consistently outperforms standard OPD and representative cross-modal distillation methods, achieving up to 6.1% and 10.7% performance gain over the strongest baselines on depth and video benchmarks, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.