Scores High Alone, Ranks Low in Context: Full-Context Knockout Frame Selection with an MLLM for Long-Video Question Answering
Abstract
Long-video question answering asks a multimodal large language model (MLLM) to answer a question about a video that runs for tens of minutes. To make this long visual input manageable, keyframe selection reduces the video to a few dozen frames. A frame can be scored alone, by what it brings when other frames are absent, or in context, by what is lost when it is removed from the complete candidate pool. At deployment the selected frames enter the model together, motivating the second condition. We find that the highest-scoring frame alone can rank near the bottom in context because other frames already supply its evidence, and the two rankings are almost unrelated. We propose KEYSTONE, which removes one temporal unit, a short run of consecutive frames, and scores the displacement of a frozen answering model's answer distribution. The reference is determined jointly by all leave-one-out distributions. The whole chain runs on the frozen answering model, without training or an external scorer. On the Video-MME, MLVU, LongVideoBench, and LVBench benchmarks, with Qwen2.5-VL, LLaVA-OneVision, LLaVA-Video, and InternVL3.5, our KEYSTONE method outperforms existing training-free frame selectors. Code will be made publicly available upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.