Training-Free Video Paragraph Grounding with Video Sentence Localizers via Sequence Consistency
Abstract
Video Paragraph Grounding (VPG) jointly localizes the temporal segments described by all sentences in a paragraph. Although paragraph-level descriptions provide semantic context and temporal relations for disambiguating repeated or visually similar events in long videos, existing VPG models can still underperform strong video sentence localizers. Through a diagnostic analysis of existing VPG methods, we find that the sequential organization of paragraph sentences provides a simple yet effective global cue for coordinating independently grounded segments. Based on this insight, we propose Sentence-to-Paragraph Grounding (S2PG), a lightweight and training-free framework that reformulates VPG as a sequence-level candidate selection problem. S2PG first obtains multiple candidate segments and confidence scores for each sentence from an off-the-shelf Video Sentence Grounding (VSG) model, and then applies Sequence-Aware Global Optimization (SAGO) to jointly optimize sentence-level grounding confidence and paragraph-level temporal consistency. Experiments on two representative VPG benchmarks demonstrate that S2PG consistently improves strong VSG baselines and achieves state-of-the-art performance among existing VPG methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.