acceptodds
Under review as a conference paper at ICLR 2027

LAMP: Learnability- and Horizon-Aware On-Policy Distillation from Privilege Modality

Abstract

Multimodal large language models (MLLMs) often struggle with perceptual information that is difficult to recover from RGB alone. Cross-modal on-policy distillation (OPD) addresses this issue by allowing the teacher to access privilege modalities such as depth or video, but existing methods treat token-level supervision uniformly, conflating transferable reasoning with unlearnable sensor information. Therefore, we propose LAMP, a learnability- and horizon-aware privilege-modality on-policy distillation framework for MLLMs. LAMP makes a two-fold contribution: First, it performs counterfactual modality attribution to identify whether each token’s teacher advantage is grounded in the student or privilege modality, enabling a self-calibrating gate to suppress untransferable supervision. Second, it introduces horizon-conditioned divergence weighting to emphasize tokens that have greater downstream influence on the reasoning trajectory. Extensive experiments on depth-to-RGB and video-to-RGB transfer show that LAMP consistently outperforms standard OPD and representative cross-modal distillation methods, achieving up to 6.1% and 10.7% performance gain over the strongest baselines on depth and video benchmarks, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.