SwiftVMR: Toward Efficient and Effective MLLM-based Video Moment Retrieval
Abstract
When repurposed for Video Moment Retrieval (VMR), Multimodal Large Language Models (MLLMs) deliver strong cross-modal alignment yet inherit significant efficiency burdens and a representational gap. Directly applying standard generative paradigms to VMR poses four major challenges to efficient and effective inference: 1) dense video sampling triggers a catastrophic visual token explosion; 2) auto-regressive decoding introduces inherent sequential overhead; 3) discrete-token formulations introduce a representational gap when modeling the continuous, ambiguous nature of action boundaries; and 4) massive parameter scales prohibit resource-constrained deployment. To address these challenges, we propose SwiftVMR, a framework retaining the MLLM's alignment prowess while replacing iterative decoding with single-shot, continuous boundary prediction. To tackle the token explosion, a Zero-initialized Motion-aware Spatial Compressor (Z-MSC) is introduced to aggressively reduce visual tokens while preserving spatial structures and training stability. To overcome the sequential decoding latency and representational mismatch of autoregressive generation, we introduce a Distribution-Aware Perception Head (DAPH), a single-shot logit-based architecture directly modeling probability distributions over temporal proposals. Finally, to overcome the deployment barrier, we introduce a Dual-Space Structured Distillation (DSSD) strategy to distill a high-capacity teacher into a lightweight student. Leveraging the continuous logit space provided by DAPH, DSSD performs foreground-weighted feature alignment and quality-gated distribution distillation. Extensive experiments on four benchmarks demonstrate that SwiftVMR matches state-of-the-art performance while accelerating inference by an order of magnitude, proving that efficiency can be regained without sacrificing performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.