acceptodds
Under review as a conference paper at ICLR 2027

Query-Answer-Aligned Frame Selection for Long Video Understanding

Abstract

Multimodal large language models (MLLMs) process multimodal inputs by converting text, images, and videos into token sequences that are then processed by a backbone language model. However, the limited visual token budget of these models makes long-form video understanding challenging. For instance, a 5-minute video at 24 frames per second (fps) contains 7,200 frames. However, practical MLLM inference typically uses only a subset of these frames, usually ranging from 8 to 64 frames. These frames are usually sampled uniformly, regardless of their relevance to the question. Recently, training-free, model-agnostic frame-selection methods have been proposed to improve video question answering. In this work, we propose two training-free, model-agnostic frame-selection methods: QV and QAV. QV scores frames using question-frame cosine similarity. QAV additionally incorporates multiple-choice answer options and scores frames by their maximum cosine similarity across the resulting question-option pairs. Both methods first construct a candidate pool by sampling video frames at a fixed rate. They then select the highest-scoring frames, up to a predefined frame budget, and restore their chronological order before passing them to the downstream MLLM. We evaluate QV and QAV on MLVU, Video-MME, and LongVideoBench using three MLLMs: Qwen2-VL-7B, LLaVA-Video-7B, and InternVL3.5-8B. Our results show that QAV improves over uniform sampling across the evaluated model-benchmark settings and achieves competitive performance against the evaluated training-free frame-selection baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.