Short-Moment Grounding in VideoLLMs: A Search-Space Limit Rather Than a Capability Limit
Abstract
Video large language models (VideoLLMs) have become the dominant approach to temporal grounding, and their progress is measured by averages over entire benchmarks. We find that these averages hide a near-total failure on short moments, which are rare in standard test sets but often matter most in practice. The failure appears in every model and dataset we examine, and we trace it to two sources: a bias toward predicting spans of a typical length and an imprecise estimate of where the moment lies. The bias is already present before any grounding training, and both sources persist across a wide range of fine-tuning interventions. The same frozen model, however, localizes short moments precisely once its input is narrowed around them, which suggests that the failure lies in how much of the video the model must search rather than in what it can perceive. We turn this finding into a training-free temporal zoom that gives a second, focused look only to moments the model itself predicts to be short, and it substantially improves short-moment grounding. Our results show that VideoLLMs can already do more than their average scores suggest, and that the cheapest progress lies at inference time.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.