acceptodds
Under review as a conference paper at ICLR 2027

SlashVID: Training-Free Extreme Token Compression via Select-and-Aggregate for Video Temporal Grounding

Abstract

Video Temporal Grounding (VTG) using Large Vision-Language Models (LVLMs) requires dense temporal sampling to preserve boundary-sensitive evidence, yet the resulting large number of visual tokens renders inference prohibitively expensive. At extreme compression ratios, pruning-only baselines can remove all tokens from some frames and discard subtle boundary cues. We propose **SlashVID**, a training-free framework that decomposes extreme token compression into two complementary paths: aggregating redundant tokens to preserve temporal context and selecting a small set of query-relevant tokens to sharpen boundary evidence. Specifically, SlashVID employs Adaptive Segment-wise Guidance to partition each video into segments and allocate a fixed budget across the two paths, Bidirectional Temporal Aggregation Clustering to consolidate the unselected context, and Dynamic Max-Min Importance Selection to add diverse, query-guided evidence. Experiments on Charades-STA and ActivityNet show that SlashVID retains 80% of full-token grounding performance at a 2% retention ratio while achieving 9.6 prefill speedup, improving relative performance retention over the strongest baseline by an absolute 4.55%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.