Learning to Look Beyond a Single Frame for Video Understanding
Abstract
Large multimodal models (LMMs) have achieved incredible success in video understanding. However, they exhibit a fundamental limitation: single-frame bias. Instead of exploiting the full temporal sequence, they base their understanding on static cues from a few frames, typically the first or the last frame. This shortcut undermines the model’s ability to develop a holistic understanding of video content, often resulting in systematic errors. To mitigate this problem, we propose MF-GRPO, a novel framework that employs reinforcement learning (RL) to encourage LMMs to engage with the full temporal context. At its core is the multi-frame focus advantage enhancement, which is triggered using a contrastive objective. Specifically, the model receives the advantage enhancement when its understanding of the entire video sequence surpasses its understanding of a partially masked version of the same sequence. This enhancement mechanism explicitly encourages the model to integrate information across the entire temporal span, beyond the excessive reliance on superficial cues of a single frame. Furthermore, this RL training is designed to be progressive. We start with shorter video sequences and gradually extend their length, enabling the model to develop the multi-frame focus ability smoothly from easy to hard. Quantitative experiments show MF-GRPO effectively improves the model performance on various video understanding benchmarks. Critically, qualitative visualizations confirm that our model learns to allocate attention more broadly across frames, successfully mitigating temporal blindness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.