acceptodds
Under review as a conference paper at ICLR 2027

Partition and Select: Efficient Training-Free Keyframe Selection for Long-Video Understanding

Abstract

Long-video understanding is challenging because videos may contain tens of thousands of frames, making it impractical to process them entirely with multimodal large language models (MLLMs). Query-aware keyframe selection addresses this issue, but existing methods often require additional training or process large candidate sets with expensive encoders. We introduce **Part**ition and **S**elect (PartS), a training-free framework for efficient query-aware keyframe selection in long videos. Our method improves efficiency by reducing both the number of candidate frames that must be processed and the cost of encoding each candidate. Specifically, PartS exploits video-codec structure to construct a compact set of visual anchors and adaptively enriches it with frames from query-relevant segments. It then encodes candidates using a lightweight CLIP model, clusters them by visual similarity, and applies query-aware soft frame allocation to distribute a fixed frame budget across clusters according to their relevance. We evaluate PartS on four long-video QA benchmarks and four open-source MLLMs. It achieves - speedups over existing training-free approaches, maintaining competitive performance. These results demonstrate that PartS is an efficient plug-and-play module for long-video understanding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.