SAVE: Let VLMs Allocate Their Own Viewing Budgets for Long-Video Frame Selection
Abstract
Long-video understanding under a limited visual budget requires selecting a small set of frames that captures question-relevant evidence. Existing query-aware selectors often rely on external models to score frame relevance. Balancing relevance with temporal coverage remains challenging: prioritizing high-scoring frames can yield redundant views of a narrow interval, whereas broad coverage can allocate scarce frames to irrelevant content. We propose SAVE (Self-guided Allocation of Visual budgets via Evidential demand), a VLM-native framework for visual budget allocation. We use the same frozen vision-language model (VLM) for evidence scoring and final question answering, converting frame-level relevance scores into a semantic demand distribution over time. Given a fixed frame budget, we formulate selection as an optimal transport problem that minimizes the demand-weighted squared temporal distance from each candidate frame to its nearest selected frame. This objective couples relevance and temporal coverage, adaptively allocating the visual budget without a separate balancing coefficient. Exploiting the one-dimensional temporal structure, we solve this objective exactly and efficiently via dynamic programming, guaranteeing a globally optimal allocation under the proposed objective. Experiments across multiple long-video benchmarks and diverse VLM architectures show consistent improvements over the evaluated baselines in benchmark-level average performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.