acceptodds
Under review as a conference paper at ICLR 2027

Event-Adaptive Temporal Patching for Token-Budgeted Video Grounding and Chaptering

Abstract

Video multimodal large language models (MLLMs) typically compress a video into a fixed number of visual tokens before temporal reasoning, using uniform sampling and fixed temporal patches. This can underrepresent brief actions and transitions, whereas increasing the sampling rate usually increases downstream processing cost. We introduce Event-Adaptive Temporal Patching (EATP), a query-independent, training-free front end that decouples observation density from the number of encoded visual units. EATP observes a dense candidate stream and uses exact dynamic programming over coarse visual descriptors to partition it into contiguous, variable-length temporal groups under a fixed visual- token budget. The resulting schedule allocates shorter groups to regions with large descriptor changes and longer groups to relatively stable content. A length- aware temporal patch embedding, derived from the pretrained two-frame kernel, encodes each group without fine-tuning. Experiments on QVHighlights moment retrieval and VidChapters chapter generation show training-free improvements over native uniform sampling, supporting the benefit of combining dense observation with content-adaptive temporal grouping rather than dense sampling alone. On QVHighlights, EATP improves mIoU by 2.16 points, strict R@[email protected] by 8.04 points and R@[email protected] by 3.14 points. On VidChapters, it improves F1, tIoU, SODA, CIDEr, and GRACE by 4.5, 1.4, 1.7, 6.5, and 1.0 points, respectively.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.