acceptodds
Under review as a conference paper at ICLR 2027

DTP-OPSD: Reinforcement Learning with Distractor-Aware Temporal-Privilege On-Policy Self-Distillation for Temporal Video Grounding

Abstract

Temporal Video Grounding (TVG) aims to localize the start and end times of an event described by a language query in a video. The task becomes particularly challenging when the same video contains multiple repeated, similar, or opposite events. On these hard examples, all responses sampled by Group Relative Policy Optimization (GRPO) may miss the target and receive zero temporal Intersection-over-Union (tIoU). Their group-relative advantages are then all zero, leaving no useful policy gradient signal. Further analysis shows that the model is often able to recognize the queried event but is distracted by other events in the same video and selects the wrong occurrence. To address this, we propose DTP-OPSD, a reinforcement learning framework with distractor-aware temporal-privilege on-policy self-distillation. The student observes the clean video and generates responses on-policy, while a shared-weight teacher observes a privileged video in which the target window is marked by red full-frame borders and same-video distractor windows by blue full-frame borders. For the distillation loss, the teacher evaluates student-generated prefixes and provides dense token-level Jensen–Shannon divergence supervision. This supervision remains informative even when GRPO provides no policy gradient signal. We suppress distillation when the teacher does not outperform the student and strengthen it for groups where the teacher performs better but all student rollouts receive zero tIoU. We also construct a compact hard-interference training set through VLM screening and human verification. With only 300 additional post-training examples, DTP-OPSD achieves the highest mIoU among the compared open-source 8B models on Charades-TimeLens, ActivityNet-TimeLens, and QVHighlights-TimeLens, reaching 55.93, 54.77, and 66.71, respectively. Compared with Pure GRPO trained on the same data, DTP-OPSD improves all 12 reported metrics, with [email protected] gains of 1.28, 1.62, and 1.78 points on the three benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.