Budgeted Multimodal Evidence Selection for Long-Video Understanding via Knapsack Optimization
Abstract
Understanding long videos with vision–language models (VLMs) requires selecting useful information within limits on the frames and text a model can process. Existing approaches address this by leveraging retrieval augmentation to supplement sampled frames with relevant information, such as speech transcripts and on-screen text, and select the most related evidence items based on similarity thresholds and ranking cutoffs. However, these heuristics do not directly optimize the joint usefulness of heterogeneous, unequal-cost evidence under a fixed model input budget. We make the first attempt in long-video understanding to formulate the final selection of unequal-cost textual evidence as a knapsack-constrained optimization problem, and introduce RCES, a training-free framework that allocates text and frames under separate input budgets without modifying the answering VLM. Specifically, RCES optimizes a monotone submodular objective that balances question relevance, temporal and source coverage, and representativeness of the selected evidence under a knapsack constraint, which can be solved by a density-greedy algorithm with performance guarantees on the objective value. A separate frame selector is further proposed to combine uniformly spaced frames for global context with short frame sequences around question-relevant moments. Across three long-video benchmarks and five VLM backbones, RCES achieves competitive accuracy with substantial reductions in textual evidence cost or latency, compared to related baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.