GLoR: Global-Local Ranking for Video Temporal Grounding
Abstract
Video temporal grounding (VTG) requires a model to return the temporal segment in an untrimmed video that corresponds to a natural-language query. In VTG, especially in long-video scenarios, candidate segments with lower relevance may receive higher confidence scores than target segments. Consequently, the model may fail not because the correct moment is absent, but because an inferior candidate is promoted to Top-1, turning an otherwise successful localization into an incorrect prediction. To address this issue, we propose a method named the Global-Local Ranking Grounder (GLoR), which decouples global candidate ranking from local boundary localization. GLoR adopts a global branch for candidate scoring and ranking and a local branch for boundary prediction. The global branch uses relevance evidence to determine the degree of match between a candidate segment and the query, while the local branch preserves high-frequency temporal variations and obtains more precise temporal boundaries by predicting probability distributions over start and end boundaries. In addition, IoU-based ranking and candidate-quality calibration further improve the accuracy of candidate scores. We evaluate GLoR on the TaCoS, Ego4D-NLQ, and Soccer-GMR datasets. Experimental results show that GLoR achieves state-of-the-art (SOTA) performance on TaCoS, Ego4D-NLQ, and Soccer-GMR, fully demonstrating its outstanding capability in candidate ranking and precise temporal localization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.