User-Controllable Dense Video Captioning: A Large-Scale Benchmark and Framework
Abstract
Dense video captioning (DVC) aims to generate temporally localized captions for multiple events in untrimmed videos. Despite recent advances, existing methods produce captions at a fixed granularity because current benchmarks provide only single-style annotations, leaving variations in event density and caption depth largely unexplored. To address this gap, we present User-Controllable Captions (UC Captions), a new dataset with annotations that vary in event density (i.e., how frequently events are localized) and caption depth (i.e., the level of descriptive detail provided for each event). This is the first DVC dataset to explicitly encode controllable dimensions of annotation. Building on this, we propose User-Controllable DVC (UC-DVC), a framework that incorporates user-defined density and depth parameters to dynamically adjust event localization and caption generation. Direct evaluation shows that UC-DVC changes primarily along the requested control axis, produces ordered responses at unseen intermediate inputs, and preserves standard DVC performance across three backbones. These results establish UC Captions as a benchmark for evaluating controllable granularity and UC-DVC as a strong reference framework for the task. To support further research, both UC Captions and UC-DVC code will be publicly released after review.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.