acceptodds
Under review as a conference paper at ICLR 2027

Q-GPS: Query-Guided Progressive Frame Selection for Long Video Question Answering

Abstract

Frame selection enables multimodal large language models (MLLMs) to process long videos within limited visual-token budgets by retaining only a small subset of frames. Highly relevant frames may concentrate within the same short interval, leaving other query-relevant moments unrepresented. Effective selection therefore requires assessing both candidate relevance and the additional temporal coverage each frame provides given previous selections. We propose Query-Guided Progressive Frame Selection (Q-GPS), a training-free, model-agnostic method that progressively selects frames according to relevance-calibrated gains in query-weighted temporal coverage. It initializes its residual state with a query-independent path-spectral temporal covariance that models temporal coupling, and uses this state as a surrogate for the temporal coverage remaining after previous selections. At each step, it **selects** the frame with the largest relevance-calibrated coverage gain, **conditions** the covariance on the selected frame's temporal location, and **rescores** the remaining candidates using the updated residual state. Across controlled evaluations with Qwen2.5-VL-7B and frame budgets of 8, 16, and 32, Q-GPS improves average accuracy over uniform sampling by 4.70%, 7.28%, and 11.05% on Video-MME, LongVideoBench, and MLVU, respectively. Code is available at https://anonymous.4open.science/r/Q-GPS-FEB4.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.