TEMPORAL RESOLUTION IS DURATION-DEPENDENT IN GENERATIVE VIDEO GROUNDING
Abstract
Temporal resolution is usually treated as uniformly beneficial in video temporal grounding. We show instead that, for a strong generative grounder, its marginal value is signed and duration-dependent. After averaging over sampling phase, increasing density from 4 to 8 fps improves one-second events, degrades medium-duration events, and crosses zero at 2.3 s, while the aggregate effect is null. Controlled interventions narrow the source: the short-event gain does not require new visual content, since duplicating byte-identical frames reproduces and exceeds the gain from real densification, while granularity-matched permutation of the model's interleaved time markers largely attenuates it, so readable temporal addressing drives the effect. Matched fine-tunes with different duration compositions also shift the density-response crossing, consistent with a learned duration-sensitive prior. We therefore introduce Zoom-VTG, an inference-time, duration-gated coarse-to-fine policy that applies a denser second pass only when the coarse prediction is short. On held-out Charades TEST-400, it improves one-second mIoU by +0.172, and the gain transfers without retuning to ActivityNet and QVHighlights. On the complete, descriptively-accounted 3,363-query population, one-second mIoU improves by +0.137 while aggregate mIoU changes by only +0.001. These results indicate that temporal resolution should be allocated and evaluated conditionally on event duration, rather than treated as a single global hyperparameter.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.