FULCRUM: Option-Contrastive Frame Selection for Long-Video Understanding
Abstract
Video-capable vision-language models (VLMs) answer questions about long videos under a limited frame budget, making frame selection critical to the evidence available to the answering model. However, question relevance alone does not indicate whether a frame supports one candidate answer over the others. We introduce FULCRUM, a training-free frame selector for multiple-choice video question answering built on option-contrastive relevance. Using a frozen image–text encoder, FULCRUM scores each frame against each question–option pair and combines the strongest match with its margin over the mean match across options. This captures both relevance to an answer hypothesis and its support relative to competing hypotheses. The resulting scores guide budget allocation across temporal segments, while a coverage floor preserves broad temporal coverage and diversity-aware selection reduces redundancy. FULCRUM requires a single encoding pass over the candidate frames and option hypotheses, without querying the answering model during selection. Under a shared evaluation protocol on LongVideoBench and Video-MME, with Qwen2.5-VL-7B and LLaVA-Video-7B, FULCRUM outperforms uniform sampling and reimplementations of ten published selectors at all evaluated budgets of 4, 8, 16, and 32 frames. With eight frames on LongVideoBench, it improves Qwen2.5-VL-7B accuracy from 44.2% to 62.5%, exceeding the strongest competing selector by 3.1 percentage points and surpassing Qwen2.5-VL-72B with Adaptive Keyframe Sampling at the same frame budget. Across both models and benchmarks, four FULCRUM-selected frames outperform 32 uniformly sampled frames.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.