acceptodds
Under review as a conference paper at ICLR 2027

FATE: Frame-and-Token Evidence Preservation for Efficient MLLM-based Video Temporal Grounding

Abstract

Multimodal large language models (MLLMs) have shown strong capabilities for video temporal grounding (VTG), but processing video inputs incurs substantial computation and memory costs, especially as video duration increases. Existing efficiency methods typically reduce visual computation through frame selection before visual encoding or token compression after encoding, with most approaches focusing on individual reduction stages. However, these stages are closely connected in VTG: frame selection determines which temporal evidence is available for encoding, while token compression determines which fine-grained visual information remains available for localization. We introduce FATE, a training-free framework that coordinates visual reduction across these stages through shared temporal evidence. FATE determines where to look through coverage-preserving frame selection and how much to preserve through relevance-adaptive token allocation. After visual encoding, query-motion-guided token preservation further determines what to preserve within each temporal patch through informative seed selection and coverage-aware completion. Together, these components provide a consistent evidence-guided strategy for preserving grounding-relevant information across the visual processing pipeline without additional training of the grounding MLLM. Experiments on three VTG benchmarks under varying frame and visual-token budgets demonstrate consistent improvements on Ego4D-NLQ and ActivityNet Captions while maintaining competitive performance on Charades-STA.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.