acceptodds
Under review as a conference paper at ICLR 2027

Grounding All Moments in Long Videos via Retrieval, Reranking and Refinement

Abstract

Video temporal grounding aims to localize video segments that match a language query. However, existing methods typically assume a single occurrence per query, which is inadequate for long videos where the same event may appear multiple times. Existing benchmarks are also limited in this setting, as repeated valid moments are often missing from the annotations. To address this problem, we introduce AO-Bench, a human-refined benchmark for arbitrary-occurrence long video temporal grounding, where the goal is to localize all temporal spans that satisfy a query. We further propose a three-stage framework that retrieves candidate clips, reranks them, and refines their temporal boundaries. Our analysis shows that reranking is the key bottleneck, as it requires distinguishing true query-matching events from visually similar distractors. Therefore, we train the reranking module on AOVTG-66K generated by a dynamic multi-agent consensus process, in which multiple MLLM agents independently assess candidate clips, exchange feedback, and iteratively revise their decisions, while weakly supported or inconsistent predictions are filtered out to produce robust consensus labels. Experiments on AO-Bench show that ReGround achieves leading multi-occurrence grounding performance while remaining competitive in the single-occurrence setting, with lower inference cost than the multi-agent pipeline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.