Video-OPD++: Geometry-Calibrated Corrective On-Policy Distillation for Temporal Video Grounding
Abstract
On-policy distillation has recently emerged as an effective post-training paradigm for temporal video grounding, combining student-induced training states with dense teacher feedback. However, each realized Video-OPD update remains centered on the token sampled by the student, without explicitly turning teacher-supported timestamp alternatives that could improve localization into training targets. The interval structure of temporal video grounding makes such alternatives directly evaluable: with the suffix fixed, a timestamp can be replaced and the resulting interval scored by temporal IoU (tIoU). Our analysis shows that, for more than three quarters of student trajectories, the teacher support contains an alternative timestamp action that improves tIoU. Motivated by this observation, we propose Video-OPD++, a geometry-calibrated corrective on-policy distillation framework. At each student-induced timestamp state, Video-OPD++ evaluates teacher-supported candidate actions using annotation-derived tIoU utilities under fixed-suffix interventions and reweights the teacher prior to construct a corrective distribution aligned with localization quality. This distribution replaces the original local sampled-action reward at timestamp positions, while a discounted future residual preserves feedback from subsequent token discrepancies. Across three temporal video grounding benchmarks, Video-OPD++ improves average mIoU over Video-OPD by 2.3 points. These results suggest that incorporating temporal video grounding geometry into the distillation target helps the student learn more accurate temporal boundaries.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.