acceptodds
Under review as a conference paper at ICLR 2027

Zero- and Multi-Event Temporal Grounding with Multimodal Large Language Model

Abstract

Video Temporal Grounding (VTG) aims to localize video segments matching a text query. Conventional VTG methods assume that each query corresponds to exactly one segment, an assumption violated in real-world scenarios where events may occur zero, once, or multiple times. We address this gap by formalizing Any-Moment Temporal Grounding (AMTG), which requires models to predict an arbitrary number of temporal segments, including none. To facilitate research on AMTG, we construct AnyMoment dataset, which includes queries with zero, single, and multiple ground-truth segments. Recently, reinforcement learning (RL) has emerged as an effective paradigm for optimizing multimodal large language models (MLLMs) for VTG. However, AMTG poses two additional challenges: recognizing distinct occurrences of the queried event from complex video content, and providing fine-grained learning signals for individual event predictions, as conventional sequence-level RL objectives assign a single advantage to the entire response and may therefore penalize correct predictions together with incorrect ones. To address these challenges, we propose Event-wise Grounding Optimization (EGO), an RL-based framework for AMTG. First, EGO employs a bivariate difficulty-aware sampling strategy that jointly characterizes the difficulty of recognizing event occurrences and locating their temporal extents, enabling RL to focus on samples that are challenging yet learnable. Second, we propose an occurrence-aware chain-of-thought reasoning strategy that identifies fine-grained descriptions of individual event occurrences and transforms them into discriminative subqueries, facilitating the localization of multiple similar or closely related events. Third, we introduce fine-grained event-level credit assignment, which computes event-specific advantages within each rollout group and assigns them to the corresponding output tokens, explicitly reinforcing correct event predictions while penalizing erroneous ones. Experiments show that EGO achieves state-of-the-art performance on AnyMoment and public multi-segment VTG benchmarks, demonstrating its effectiveness across diverse temporal grounding scenarios.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.