A Dynamic Temporal Segmentation for Long Video Understanding
Abstract
Multimodal Large Language Models (MLLMs) have achieved remarkable success in video understanding, yet scaling them to long videos remains a major challenge due to severe temporal redundancy and limited frame budgets. Some existing methods rely on uniform temporal segmentation, implicitly assuming that semantic content varies uniformly over time, which misallocates frame budgets and misses critical moments. In this paper, we propose DyQCA, a Dynamic Segmentation framework for keyframe selection in long video understanding. DyQCA detects temporal boundaries by comparing the semantic differences of adjacent frames, partitions the video into dynamic segments, and allocates the frame budget across segments by jointly modeling query relevance and content deviation, followed by anchor-centric keyframe selection within each segment. The proposed DyQCA is training-free, model-agnostic, and can be seamlessly plugged into existing Video-LLMs. Extensive experiments demonstrate that DyQCA consistently outperforms uniform sampling and recent keyframe selection baselines, e.g., 67.3% on LongVideoBench and 76.2% on MLVU under a 64-frame budget, demonstrating that dynamic semantic segmentation offers a simple yet effective path toward more accurate long video understanding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.