VerifyGround: Learning Verification-Guided Evidence Acquisition for Video Temporal Grounding
Abstract
Video temporal grounding requires visual evidence that identifies the queried event and resolves its temporal boundaries. Many existing multimodal large language model (MLLM)-based approaches predict temporal intervals from fixed global observations. Although these observations provide broad temporal coverage, they may be insufficient for accurate event retrieval and boundary refinement. We introduce VerifyGround, an interactive temporal grounding framework in which verification guides subsequent observation. The model first proposes a candidate interval, then inspects the corresponding video segment to verify its event match and boundary support. It acquires additional visual evidence from alternative temporal regions when the event hypothesis lacks support and gathers denser boundary evidence when the event is supported but its temporal extent remains uncertain. Verification thus guides prediction correction through targeted evidence acquisition, determining where and how to observe next. To learn this policy, we construct GroundStep-30K for cold-start supervised fine-tuning and introduce process-aware reinforcement learning to address the difficulty of assigning credit to intermediate corrections from final outcomes alone. Training combines the final localization objective with action-specific progress supervision, crediting each decision turn for the improvement produced by its corrective action. VerifyGround achieves state-of-the-art temporal grounding performance on Charades, ActivityNet, and QVHighlights, with a 3.4 points mIoU improvement over TimeLens-8B on QVHighlights. It also improves mIoU by 1.32 points over the same baseline on TACoS under cross-dataset zero-shot evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.