TIDE: Temporal Information-Dense Evidence Compression for Video Temporal Grounding
Abstract
Video temporal grounding (VTG) usually requires Large Vision-Language Models (LVLMs) to localize query-relevant events and accurately predict their start and end timestamps in untrimmed videos, incurring substantial computational and memory costs. Existing token compression methods, mostly developed for general video understanding, reduce visual redundancy by pruning or merging tokens based on feature similarity or visual saliency. However, these appearance-based criteria are not well aligned with temporal grounding: precise localization relies on subtle state transitions and coherent entity evolutions across frames rather than visual appearance alone. Pruning or merging tokens based solely on visual appearance can fragment entity trajectories and attenuate temporal contrasts between adjacent states. To address this problem, we propose TIDE, a token compression framework that shifts the objective from removing appearance-based redundancy to preserving temporal evidence and maintaining structural continuity. TIDE identifies piecewise-stable states by tracking cross-frame token correspondences, allocates token budgets based on query-conditioned transition signals, and recovers peripheral context through soft state-compatible routing. At a visual token retention rate of 12.5% on ActivityNet-Captions, TIDE outperforms the full-token baseline in mIoU with Qwen3-VL-8B and achieves a 4.7 speedup for compression plus LLM prefilling with Qwen3-VL-4B. The code is available at https://anonymous.4open.science/r/TIDE-1125.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.