CoverVTG: Coverage-Constrained Sparse Observation for Universal Video Temporal Grounding
Abstract
Universal Video Temporal Grounding (VTG) aims to localize diverse natural-language queries in videos spanning different domains, viewpoints, and durations. Recent approaches largely rely on multimodal large language models (MLLMs), whose large parameter scale and dense visual encoding cost make them expensive for long videos and difficult to deploy in resource-constrained scenarios. In this work, we explore a lightweight backbone-centric alternative and propose CoverVTG, a coverage-constrained framework that directly reduces the number of visual tokens requiring expensive backbone encoding. Since shallow patch embeddings lack reliable high-level semantics, CoverVTG formulates pre-backbone sparsification as a structured sparse observation problem rather than semantic token selection. Within each local spatio-temporal tube, a sparse set of tokens is selected under explicit spatial and temporal coverage constraints, so that discarded observations remain locally observable from complementary retained tokens. We formalize this property through a coverage–observability–recoverability connection and establish a bound on local recovery error under spatio-temporal smoothness. Lightweight tube-local recovery adapters inserted throughout a frozen visual backbone progressively exploit such complementary observations as semantic representations emerge. Extensive experiments show that CoverVTG achieves competitive grounding accuracy with substantially fewer parameters and lower inference latency than MLLM-based approaches.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.