Temporal Quantization for Adaptive Candidate Frame Construction in Video QA
Abstract
Video language models (VLMs) operate under limited visual-token and computational budgets, requiring videos to be represented by a compact set of frames. A common pipeline first constructs a temporally subsampled candidate frame set and then selects the final input frames from it. While conventional pipelines rely on query-agnostic temporal sampling, recent work has increasingly explored adaptive frame acquisition guided by the query and partial observations. Yet, how a fixed observation budget should represent the video timeline before downstream selection remains less explicitly characterized. To address this, we introduce QUEST, a training-free method that formulates candidate construction as relevance-weighted temporal quantization, explicitly measuring how well the acquired observations resolve the timeline. By jointly considering estimated query relevance and temporal gaps between observations, QUEST allocates finer temporal resolution to relevant yet under-resolved regions while maintaining coverage elsewhere. The resulting candidate set can be used with existing frame selectors without modifying their selection rules. Extensive experiments across multiple VLMs, frame selectors, and video understanding benchmarks demonstrate improved downstream accuracy under matched observation and selection budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.