acceptodds
Under review as a conference paper at ICLR 2027

Through the LENS of Learnability: Privileged Self-Distillation for Video Reasoning

Abstract

On-policy self-distillation (OPSD) has emerged as an effective post-training paradigm for video large language models (Video-LLMs), where a privileged self-teacher provides token-level guidance along student-generated trajectories. For video understanding, however, privileged information can take diverse forms, including answer hints, textual descriptions, and additional temporal observations, whose learning value is not captured by teacher accuracy alone. Through systematic analysis, we find that stronger teachers do not necessarily produce better students; whether the privileged evidence is actually learnable by the student is the key factor. We therefore revisit video OPSD through the lens of privilege learnability, governed by two factors: teacher reliability and student visibility. Based on this analysis, we propose LENS, a learnable privileged evidence framework for video OPSD. LENS constructs privileged teachers from densely grounded temporal evidence, estimates teacher reliability at the sample level, filters supervision beyond the student's observable scope, and routes each sample to distillation or reinforcement supervision based on teacher reliability. Extensive experiments across five challenging video reasoning benchmarks demonstrate that LENS consistently outperforms existing approaches, with the distilled student even surpassing its privileged teacher on temporal grounding tasks, validating the importance of learnability-aware privileged distillation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.