COMPASS: Structured Evidence Selection for Long-Video Understanding
Abstract
Long-video understanding with Multimodal Large Language Models (MLLMs) remains challenging because raw video sequences contain far more visual information than can be accommodated within limited frame or token budgets. Existing question-driven frame selection methods typically rank frames independently based on query-frame similarity or temporal statistics, often selecting redundant frames, overlooking critical transitions, and failing to allocate visual evidence effectively across temporally complex videos. We propose COMPASS (Content-aware Multi-granularity Perception and Adaptive Semantic Selection), a deterministic and training-free framework for question-driven long-video understanding. COMPASS treats keyframe selection as a structured evidence allocation problem by jointly modeling query intent, video content structure, and temporal granularity. It derives the temporal scope and information requirements of a query, identifies coherent and informative regions, and performs set-level selection to balance query relevance, semantic coverage, and temporal continuity. A hierarchical allocation strategy further concentrates frame budgets on salient temporal windows while preserving global context. We evaluate COMPASS on LongVideoBench, Video-MME, LVBench, and MLVU under matched frame budgets. Across diverse visual encoders and downstream MLLMs, COMPASS consistently improves reasoning accuracy over existing selection strategies. On LongVideoBench with Qwen3-VL-8B, it achieves a 10.4% relative improvement over uniform sampling under identical frame budgets, with only 8.6 ms of CPU selection latency. These results demonstrate the effectiveness of structured, query-aware evidence allocation for efficient long-video understanding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.