acceptodds
Under review as a conference paper at ICLR 2027

Guide the Tokens, Rank the Moments: Adaptive On-Policy Distillation for Temporal Video Grounding

Abstract

Temporal video grounding requires identifying a queried event and recovering its temporal extent: an answer can contain the right action yet include irrelevant context or omit part of the event. On-policy distillation supplies dense teacher guidance during timestamp generation, but matching token distributions does not explicitly require the student’s preferences over complete answers to reflect their localization quality. Our central idea is to use teacher proposals both as answers to imitate and as alternatives whose relative quality can be evaluated through temporal overlap. We propose Geo-OPD, a geometry-aware on-policy distillation framework that combines pointwise imitation of a teacher answer with listwise supervision, aligning the student’s relative sequence scores with an annotation-IoU-derived preference distribution over teacher proposals. Adaptive token guidance complements these sequence objectives by balancing soft and sharp views of the same frozen teacher according to temporal discrepancy, while reliability filtering selects teacher supervision. Experiments on Charades-TimeLens, ActivityNet-TimeLens, and QVHighlights-TimeLens show improvements over the reported distillation baselines, with component ablations indicating that the additional benefit of listwise supervision varies across datasets. These results support combining guidance on how to generate temporal answers with explicit supervision on which candidate intervals better localize the queried event.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.