OPD-TV: Rethinking Teacher Views for Privileged Supervision Transfer in Video On-Policy Distillation
Abstract
Privileged on-policy distillation extends on-policy distillation by allowing a teacher model to construct supervision from temporal evidence unavailable to the student. In Video OPD, this privilege is instantiated through additional temporal evidence available to the teacher while the student continues to operate under its original visual input. We find a supervision transfer gap: continuously enriching the teacher’s visual view improves the teacher’s grounding accuracy yet produces non-monotonic student gains. This supervision transfer gap arises because an enriched teacher can exploit visual cues unavailable to the student, introducing target shifts that the student cannot effectively absorb. To address this mismatch, we introduce OPD-TV, which regulates the transfer of privileged information based on token-level cross-view compatibility. By comparing the same teacher across both visual states, the proposed method explicitly isolates the additional corrections introduced by changing the teacher view from the standard policy correction in same-view distillation. It then anchors the transfer on the same-view teacher distribution and dynamically scales the transfer strength according to how closely the enriched teacher distribution remains aligned with the student-view teacher distribution. This continuous regulation selectively transfers useful visual refinements while suppressing corrections that are misaligned with the student’s visual input. Across three temporal grounding benchmarks, OPD-TV consistently improves privileged supervision transfer.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.