A Bimodal Perspective on Video Grounding
Abstract
Video temporal grounding requires localizing events within videos based on natural language queries. Although temporal Intersection over Union (tIoU) is the de facto metric for both evaluation and Reinforcement Learning from Verifiable Rewards (RLVR), we show that applying it uniformly across event durations creates two previously underexplored failure modes. First, for short events, small absolute boundary shifts cause disproportionately large changes in tIoU, so temporally close predictions can receive misleadingly low scores. Second, when tIoU is directly optimized as an RLVR reward, a model can increase its expected reward by inflating predicted intervals; consequently, tIoU-based RLVR improves long-event localization while degrading short-event performance. These observations motivate a duration-aware framework that treats short and long events as distinct temporal regimes. We evaluate them separately with distance-based and overlap-based criteria, respectively. For training, we combine distance-IoU with boundary accuracy to provide smoother supervision for short events while preserving effective overlap supervision for long events and discouraging interval inflation. We further introduce a multi-rollout query sampler that uses duration-aware success criteria to retain hard-yet-solvable examples, together with BiVTG-100K, a refined dataset with improved coverage and annotation quality for short events. Experiments on the manually re-annotated Charades, ActivityNet, and QVHighlights benchmarks show that the proposed framework consistently improves duration-aware grounding and mitigates the short-event failures of tIoU-based RLVR.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.