acceptodds
Under review as a conference paper at ICLR 2027

SEEK: Preserving Sparse Evidence across Semantic Events for Long-Video Understanding

Abstract

Long-video understanding challenges Multimodal Large Language Models (MLLMs) to identify sparse yet decisive evidence from temporal contexts under limited visual budgets. Keyframe selection provides an efficient way to reduce the visual input while preserving task-relevant information. However, existing methods may miss sparse query-relevant evidence, over-select redundant high-relevance frames, or struggle to handle heterogeneous events with different durations, relevance distributions, and evidence densities, leading to less reliable event-level keyframe allocation. We propose SEEK, an event-centric keyframe selection framework for preserving sparse evidence across semantic events. SEEK first partitions a video into variable-length events based on local and contextual visual changes. It then evaluates each event by jointly considering peak relevance, top- mean relevance, and average relevance, and employs the resulting utility to guide event-wise keyframe allocation. Finally, relevance entropy is used within each event to adapt the candidate pool and balance query relevance, visual diversity, and temporal diversity. Experiments on LongVideoBench, Video-MME, and MLVU with LLaVA-Video-7B, InternVL3.5-8B, and Qwen3-VL-8B show that SEEK consistently outperforms representative frame selection methods, with improvements of – over uniform temporal sampling.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.