TROVE: LEARNING TO RE-OBSERVE VIDEOS FOR TEMPORAL GROUNDING
Abstract
Video temporal grounding requires global context to identify a queried occurrence and fine-grained local evidence to resolve its boundaries. Uniform frame sampling provides temporal coverage, but distributes additional observations across both relevant and irrelevant regions. We propose TROVE, a framework that learns temporal grounding through targeted re-observation. A shared policy first predicts a coarse interval from sparse global frames, then refines it using dense local evidence selected by that interval together with sparse global support. We train this trajectory through Grounding Initialization followed by Joint Re-Observation Learning. A Geometry-to-Metric Reward Curriculum supplies graded feedback for disjoint predictions before emphasizing overlap quality and evaluation thresh- olds. Contrastive Re-Observation Calibration (CRC) trains the first action using both its localization reward and the difference between targeted and uniform second-observation rewards. The uniform comparator is used only during training. On TimeLens-Bench, TROVE with a 4B backbone improves mIoU over the native model by 11.7, 8.5, and 4.5 points on the Charades, ActivityNet, and QVHighlights subsets, respectively. Experiments across three model scales, training and observation ablations, and boundary analyses demonstrate the value of learning temporal localization together with targeted evidence acquisition.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.