acceptodds
Under review as a conference paper at ICLR 2027

SkimVTG: Video Temporal Grounding via Progressive Boundary Refinement

Abstract

Video Temporal Grounding (VTG) aims to localize the temporal segment in an untrimmed video that corresponds to a natural language query. Despite recent progress, precise boundary localization remains challenging because clip-query relevance scores often vary smoothly around the target moment, providing limited cues for identifying exact start and end boundaries. To address this, we propose SkimVTG, a progressive boundary refinement framework that first localizes the target moment from a temporally compressed representation and then refines its boundaries toward the original temporal resolution. At each temporal resolution, Boundary Cross-Attention (BCA) combines clip–query relevance with its temporal difference to extract boundary-sensitive cues, while the Temporal Bridging Network (TBN) propagates refined coarse-level representations to finer temporal resolutions through Temporal Distance Cross-Attention (TDCA). By alternating BCA and TBN, SkimVTG progressively sharpens coarse localization into precise temporal boundaries. Experiments on QVHighlights, Charades-STA, and TACoS demonstrate strong performance in Moment Retrieval and Highlight Detection. The code will be released upon publication.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.