acceptodds
Under review as a conference paper at ICLR 2027

CoverVTG: Coverage-Constrained Sparse Observation for Universal Video Temporal Grounding

Abstract

Universal Video Temporal Grounding (VTG) aims to localize diverse natural-language queries in videos spanning different domains, viewpoints, and durations. Recent approaches largely rely on multimodal large language models (MLLMs), whose large parameter scale and dense visual encoding cost make them expensive for long videos and difficult to deploy in resource-constrained scenarios. In this work, we explore a lightweight backbone-centric alternative and propose CoverVTG, a coverage-constrained framework that directly reduces the number of visual tokens requiring expensive backbone encoding. Since shallow patch embeddings lack reliable high-level semantics, CoverVTG formulates pre-backbone sparsification as a structured sparse observation problem rather than semantic token selection. Within each local spatio-temporal tube, a sparse set of tokens is selected under explicit spatial and temporal coverage constraints, so that discarded observations remain locally observable from complementary retained tokens. We formalize this property through a coverage–observability–recoverability connection and establish a bound on local recovery error under spatio-temporal smoothness. Lightweight tube-local recovery adapters inserted throughout a frozen visual backbone progressively exploit such complementary observations as semantic representations emerge. Extensive experiments show that CoverVTG achieves competitive grounding accuracy with substantially fewer parameters and lower inference latency than MLLM-based approaches.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.