Learning to Ground Without Forgetting: Reinforcement Learning for Long-Video Temporal Understanding in Multimodal LLMs
Abstract
Long-video temporal grounding introduces complexities not found in standard short-video settings, as the extended temporal horizon amplifies localization error and increases event sparsity. A central challenge in adopting reinforcement learning (RL) for this task is preserving pre-trained knowledge while actively exploring novel strategies for generating the precise, structured timestamps needed to resolve event sparsity and ambiguity. To address this, we propose Token-aware KL Regularization, which selectively relaxes the KL-divergence on timestamp-related tokens to encourage exploration during early training. To further mitigate the sparsity of key events, we introduce a dense reward called Center Distance Reward (CenDist). Combined with our automatically constructed SceneTG dataset that facilitates effective RL training, our model, LongVTG-R1, achieves the highest mIoU among comparable-scale methods across three long-video temporal grounding benchmarks. Beyond grounding, our model generalizes to long-video question answering (QA), serving either as an explicit grounding module or, more importantly, as a unified system that jointly performs grounding and QA. Finally, despite being trained only on single-segment data, LongVTG-R1 exhibits an emergent capability for recursive multi-segment grounding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.