Beyond Static Cues: Motion-Aware Video Retrieval
Abstract
We identify a fundamental limitation in current video retrieval models: they rely heavily on static appearance cues and inadequately capture the motion dynamics that distinguish visually similar actions. To systematically study this limitation, we introduce MAVR, a motion-centric benchmark and training dataset built largely from newly collected videos for fine-grained text-to-video and video-to-text retrieval. MAVR combines motion-defined actions, scene-matched hard negatives, and a controlled caption scheme that separates action information from scene and viewpoint cues. Despite performing strongly on a broader multimodal embedding benchmark, existing models perform poorly on MAVR, revealing a substantial gap in motion-sensitive retrieval. As a first step toward closing this gap, we introduce MAVE, a motion-aware embedding architecture that augments a multimodal embedding model with temporally informed tokens from a dedicated video pathway. When trained on MAVR, MAVE substantially outperforms state-of-the-art retrieval baselines on actions represented during training, demonstrating the benefit of explicit temporal modeling. However, these gains do not consistently generalize to unseen actions underscoring that robust motion-aware retrieval remains an open challenge.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.