What is the Event and When does it Happen?
Abstract
Video Moment Retrieval fundamentally demands mapping a continuous visual stream onto discrete temporal intervals. This distinctive characteristic requires a discriminative latent space capable of determining what event is being described, when it occurs, and which candidate is most plausible. However, existing approaches remain largely optimized through localization-oriented objectives that supervise final temporal predictions, leaving the latent space only indirectly constrained with respect to such discrimination. In this paper, we explicitly decompose multi-dimensional discriminability into a triad of core axes: semantic, temporal, and relational discrimination. Building on this view, we introduce M-R1, a reinforcement learning-based framework that rethinks moment retrieval as the joint resolution of semantic, temporal, and ranking ambiguities. Specifically, we first propose MM-GRPO, which encourages the model to identify modality-specific semantic evidence and selectively regulate its participation in cross-modal interactions. We further devise TA-MOD, which structures temporal information flow within the fused representation via query-conditioned affinity modulation to preserve temporal discriminability. Finally, we present Rank-GRPO, which models inter-candidate relations and optimizes the relative ordering of competing moments. Experimental results on large-scale VMR benchmarks demonstrate significant improvements, with M-R1 achieving average performance gains of 5.4% on multi-moment retrieval and 32.4% on short-moment localization over previous state-of-the-art methods, showing its robustness in challenging scenarios.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.