acceptodds
Under review as a conference paper at ICLR 2027

VerifyGround: Learning Verification-Guided Evidence Acquisition for Video Temporal Grounding

Abstract

Video temporal grounding requires visual evidence that identifies the queried event and resolves its temporal boundaries. Many existing multimodal large language model (MLLM)-based approaches predict temporal intervals from fixed global observations. Although these observations provide broad temporal coverage, they may be insufficient for accurate event retrieval and boundary refinement. We introduce VerifyGround, an interactive temporal grounding framework in which verification guides subsequent observation. The model first proposes a candidate interval, then inspects the corresponding video segment to verify its event match and boundary support. It acquires additional visual evidence from alternative temporal regions when the event hypothesis lacks support and gathers denser boundary evidence when the event is supported but its temporal extent remains uncertain. Verification thus guides prediction correction through targeted evidence acquisition, determining where and how to observe next. To learn this policy, we construct GroundStep-30K for cold-start supervised fine-tuning and introduce process-aware reinforcement learning to address the difficulty of assigning credit to intermediate corrections from final outcomes alone. Training combines the final localization objective with action-specific progress supervision, crediting each decision turn for the improvement produced by its corrective action. VerifyGround achieves state-of-the-art temporal grounding performance on Charades, ActivityNet, and QVHighlights, with a 3.4 points mIoU improvement over TimeLens-8B on QVHighlights. It also improves mIoU by 1.32 points over the same baseline on TACoS under cross-dataset zero-shot evaluation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.