acceptodds
Under review as a conference paper at ICLR 2027

SAVE: Let VLMs Allocate Their Own Viewing Budgets for Long-Video Frame Selection

Abstract

Long-video understanding under a limited visual budget requires selecting a small set of frames that captures question-relevant evidence. Existing query-aware selectors often rely on external models to score frame relevance. Balancing relevance with temporal coverage remains challenging: prioritizing high-scoring frames can yield redundant views of a narrow interval, whereas broad coverage can allocate scarce frames to irrelevant content. We propose SAVE (Self-guided Allocation of Visual budgets via Evidential demand), a VLM-native framework for visual budget allocation. We use the same frozen vision-language model (VLM) for evidence scoring and final question answering, converting frame-level relevance scores into a semantic demand distribution over time. Given a fixed frame budget, we formulate selection as an optimal transport problem that minimizes the demand-weighted squared temporal distance from each candidate frame to its nearest selected frame. This objective couples relevance and temporal coverage, adaptively allocating the visual budget without a separate balancing coefficient. Exploiting the one-dimensional temporal structure, we solve this objective exactly and efficiently via dynamic programming, guaranteeing a globally optimal allocation under the proposed objective. Experiments across multiple long-video benchmarks and diverse VLM architectures show consistent improvements over the evaluated baselines in benchmark-level average performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.