acceptodds
Under review as a conference paper at ICLR 2027

Redirecting Teacher Supervision: Temporal-Response-Guided Latent Distillation for Video-Language Models

Abstract

On-policy distillation (OPD) transfers capabilities from large models to smaller ones through teacher supervision on student-generated trajectories. GKD and MiniLLM develop this paradigm for language models, while Video-OPD extends it to video tasks. Hidden-state distillation moves beyond output distribution matching, but direct alignment with teacher representations does not explicitly use the student's temporal responses to guide supervision. We introduce Temporal-Response Guidance (TRG), which uses the student's responses to local temporal changes to adjust the direction of teacher supervision. TRG adds the projection of the teacher correction onto the student's temporal response subspace back to the original correction without separately rescaling the projection. This preserves complementary teacher information outside the response subspace. Under an orthogonal projection and a detached target, TRG changes the supervision direction while preserving the cosine-loss value and the local gradient norm with respect to the projected student representation. Across four video understanding benchmarks—TempCompass, Video-MME, MVBench, and Video-MMMU—TRG improves average accuracy by 2.86 percentage points over the reported output-only OPD baseline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.