UMVR: A Unified and Efficient MLLM Framework for Short- and Long-Video Retrieval
Abstract
Text-based video retrieval has largely been studied under separate paradigms for short and long videos, despite sharing the same objective of measuring text-video relevance. A key obstacle to unifying these paradigms is that representation granularity is typically predefined for each retrieval setting rather than adapted to the temporal structure of individual videos. To address this limitation, we propose Length-Adaptive Event Representation (LAER), which adaptively presents each video into temporally coherent events whose number and temporal extent adapt to its content. This event-centric representation supports both localized and holistic relevance modeling, providing a common basis for short- and long-video retrieval. Building on LAER, we introduce UMVR, a unified MLLM-based retrieval framework that uses a shared backbone and scoring formulation across both settings. However, encoding videos with an MLLM incurs substantial computational overhead as the number of visual tokens grows, particularly for long videos. To improve encoding efficiency, we further introduce Event-Aware Token Reduction (EATR), which uses the event structure provided by LAER to allocate visual tokens according to event informativeness, retaining more tokens for informative events while compressing redundant ones more aggressively. Extensive experiments across seven short- and long-video retrieval benchmarks demonstrate the effectiveness and efficiency of UMVR, which maintains strong retrieval performance while reducing the visual-token load for MLLM encoding by 90%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.