FATE: Frame-and-Token Evidence Preservation for Efficient MLLM-based Video Temporal Grounding
Abstract
Multimodal large language models (MLLMs) have shown strong capabilities for video temporal grounding (VTG), but processing video inputs incurs substantial computation and memory costs, especially as video duration increases. Existing efficiency methods typically reduce visual computation through frame selection before visual encoding or token compression after encoding, with most approaches focusing on individual reduction stages. However, these stages are closely connected in VTG: frame selection determines which temporal evidence is available for encoding, while token compression determines which fine-grained visual information remains available for localization. We introduce FATE, a training-free framework that coordinates visual reduction across these stages through shared temporal evidence. FATE determines where to look through coverage-preserving frame selection and how much to preserve through relevance-adaptive token allocation. After visual encoding, query-motion-guided token preservation further determines what to preserve within each temporal patch through informative seed selection and coverage-aware completion. Together, these components provide a consistent evidence-guided strategy for preserving grounding-relevant information across the visual processing pipeline without additional training of the grounding MLLM. Experiments on three VTG benchmarks under varying frame and visual-token budgets demonstrate consistent improvements on Ego4D-NLQ and ActivityNet Captions while maintaining competitive performance on Charades-STA.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.